Integration Testing
Overview
Telemetry AI includes 19 end-to-end integration test scenarios that exercise the full stack: monitored apps generate telemetry under controlled failure conditions, the AI module analyzes it, and an LLM-as-judge scores the analysis quality (threshold: 70/100).
Running Tests
Full Test Suite
./run-integration-test.sh <chaos|db|weather> [ai-profile] [scorer]
Examples:
./run-integration-test.sh chaos # chaos, both openai
./run-integration-test.sh chaos grok # AI=grok, scorer=openai
./run-integration-test.sh chaos openai grok # AI=openai, scorer=grok
./run-integration-test.sh db grok grok # DB tests, both grok
./run-integration-test.sh weather openai grok # weather, AI=openai, scorer=grok
All Provider Combinations
./run-all-integration-tests.sh chaos # all 4 combos for chaos
./run-all-integration-tests.sh db # all 4 combos for db
./run-all-integration-tests.sh weather # all 4 combos for weather
Specific Test Methods
./run-integration-methods.sh <chaos|db|weather> <ai-profile> [scorer] <method1> [method2] ...
Pass an empty string '' for scorer to use the default (openai).
./run-integration-methods.sh chaos openai grok analyzeLatency analyzeDeadlock
./run-integration-methods.sh db grok '' analyzeSlowQueries
./run-integration-methods.sh weather openai grok analyzeWeatherSlowApi
Chaos Test Scenarios
| # | Method | What It Tests |
|---|---|---|
1 |
|
Healthy requests — no false positives |
2 |
|
HTTP 4xx/5xx — error pattern detection |
3 |
|
Thread.sleep delays — latency spike detection |
4 |
|
Memory + CPU + leaks — resource concern detection |
5 |
|
Exception + threadpool + GC — multi-chaos correlation |
6 |
|
Synchronized lock contention — thread blocking |
7 |
|
~60% random failure rate — flaky error detection |
8 |
|
App stopped/restarted — outage detection |
9 |
|
Multiple delayed requests + error — latency detection |
10 |
|
Thread deadlock — deadlock detection |
11 |
|
Dev MCP source analysis — code correlation |
12 |
|
Grafana dashboard JSON — visualization quality |
DB Test Scenarios
DB tests use Toxiproxy for infrastructure-level chaos injection:
flowchart LR App[Ext App] -->|JDBC| Toxi[Toxiproxy :33061] Toxi -->|JDBC| DB[(MySQL :33060)] style Toxi fill:#fdd,stroke:#c33
| # | Method | What It Tests |
|---|---|---|
13 |
|
Database queries — healthy state confirmation |
14 |
|
Toxiproxy 3s latency — DB slowness detection |
15 |
|
4s latency + small pool — mixed success/failure pattern |
16 |
|
Toxiproxy connection cut — outage detection |
Weather Test Scenarios
Weather tests use Toxiproxy to inject chaos on the external Open-Meteo weather API connection:
flowchart LR App[Ext App] -->|REST| Toxi[Toxiproxy :8888] Toxi -->|REST| API[api.open-meteo.com] style Toxi fill:#fdd,stroke:#c33
| # | Method | What It Tests |
|---|---|---|
17 |
|
REST client calls to weather API — healthy state |
18 |
|
Toxiproxy 3s latency on weather API — external latency detection |
19 |
|
Toxiproxy connection cut to weather API — outage and recovery |
Chaos Types
The app module supports 11 chaos types via GET /chaos?type={type}&intensity={value}:
| Type | Effect |
|---|---|
|
Thread.sleep (configurable ms) |
|
Allocate N MB (released after request) |
|
CPU burn loop for N ms |
|
Allocate N MB (never freed) |
|
Random 5xx WebApplicationException |
|
Unhandled RuntimeException (HTTP 500) |
|
Block 10 threads via CountDownLatch |
|
10 threads competing for single synchronized lock |
|
Rapid alloc/dealloc loop causing GC pressure |
|
Random failures at configurable rate |
|
Two threads deadlocked on competing locks |
Evaluation Pipeline
Each test uses an LLM-as-judge approach: a scorer LLM evaluates the analysis produced by the AI provider LLM.
sequenceDiagram participant Test as Test Case participant App as Companion App participant AI as AI Module participant Scorer as Scorer LLM Test->>App: Inject chaos Note over Test: Wait for telemetry ingestion Test->>AI: GET /analyze/n AI-->>Test: Analysis + captured tool outputs Test->>Scorer: Evaluate analysis against telemetry data Scorer-->>Test: Score (0-100) + prompt suggestions Note over Test: Assert score >= 70
The scorer evaluates on four dimensions (each 0-25 points): completeness, accuracy, correlation quality, and actionability. If the score falls below 70, the scorer also outputs concrete prompt improvement suggestions for the system prompt.
Cross-LLM Combinations
The analyzer LLM and scorer LLM can be different providers, preventing self-evaluation bias:
./run-integration-test.sh chaos openai grok # AI=OpenAI, scorer=Grok
./run-integration-test.sh chaos grok openai # AI=Grok, scorer=OpenAI
./run-all-integration-tests.sh chaos # all 4 combinations
Environment Variables
Required for integration tests:
export OPENAI_API_KEY=sk-... # required for openai profile
export GROK_API_KEY=xai-... # required for grok profile
Optional scorer overrides:
export SCORER_BASE_URL=https://api.x.ai/v1
export SCORER_API_KEY=$GROK_API_KEY
export SCORER_MODEL=grok-3-mini