Integration Testing

Overview

Telemetry AI includes 19 end-to-end integration test scenarios that exercise the full stack: monitored apps generate telemetry under controlled failure conditions, the AI module analyzes it, and an LLM-as-judge scores the analysis quality (threshold: 70/100).

Running Tests

Full Test Suite

./run-integration-test.sh <chaos|db|weather> [ai-profile] [scorer]

Examples:

./run-integration-test.sh chaos                    # chaos, both openai
./run-integration-test.sh chaos grok               # AI=grok, scorer=openai
./run-integration-test.sh chaos openai grok        # AI=openai, scorer=grok
./run-integration-test.sh db grok grok             # DB tests, both grok
./run-integration-test.sh weather openai grok      # weather, AI=openai, scorer=grok

All Provider Combinations

./run-all-integration-tests.sh chaos    # all 4 combos for chaos
./run-all-integration-tests.sh db       # all 4 combos for db
./run-all-integration-tests.sh weather  # all 4 combos for weather

Specific Test Methods

./run-integration-methods.sh <chaos|db|weather> <ai-profile> [scorer] <method1> [method2] ...

Pass an empty string '' for scorer to use the default (openai).

./run-integration-methods.sh chaos openai grok analyzeLatency analyzeDeadlock
./run-integration-methods.sh db grok '' analyzeSlowQueries
./run-integration-methods.sh weather openai grok analyzeWeatherSlowApi

Chaos Test Scenarios

# Method What It Tests

1

analyzeNormalTraffic

Healthy requests — no false positives

2

analyzeErrorTraffic

HTTP 4xx/5xx — error pattern detection

3

analyzeLatency

Thread.sleep delays — latency spike detection

4

analyzeResourcePressure

Memory + CPU + leaks — resource concern detection

5

analyzeCascadingFailure

Exception + threadpool + GC — multi-chaos correlation

6

analyzeLockContention

Synchronized lock contention — thread blocking

7

analyzeIntermittentFailures

~60% random failure rate — flaky error detection

8

analyzeNetworkPartition

App stopped/restarted — outage detection

9

analyzeRequestFlood

Multiple delayed requests + error — latency detection

10

analyzeDeadlock

Thread deadlock — deadlock detection

11

examineSourceCode

Dev MCP source analysis — code correlation

12

generateDashboard

Grafana dashboard JSON — visualization quality

DB Test Scenarios

DB tests use Toxiproxy for infrastructure-level chaos injection:

flowchart LR
App[Ext App] -->|JDBC| Toxi[Toxiproxy :33061]
Toxi -->|JDBC| DB[(MySQL :33060)]
style Toxi fill:#fdd,stroke:#c33
# Method What It Tests

13

analyzeNormalTraffic

Database queries — healthy state confirmation

14

analyzeSlowQueries

Toxiproxy 3s latency — DB slowness detection

15

analyzePoolExhaustion

4s latency + small pool — mixed success/failure pattern

16

analyzeDbOutage

Toxiproxy connection cut — outage detection

Weather Test Scenarios

Weather tests use Toxiproxy to inject chaos on the external Open-Meteo weather API connection:

flowchart LR
App[Ext App] -->|REST| Toxi[Toxiproxy :8888]
Toxi -->|REST| API[api.open-meteo.com]
style Toxi fill:#fdd,stroke:#c33
# Method What It Tests

17

analyzeWeatherNormal

REST client calls to weather API — healthy state

18

analyzeWeatherSlowApi

Toxiproxy 3s latency on weather API — external latency detection

19

analyzeWeatherApiOutage

Toxiproxy connection cut to weather API — outage and recovery

Chaos Types

The app module supports 11 chaos types via GET /chaos?type={type}&intensity={value}:

Type Effect

delay

Thread.sleep (configurable ms)

memory

Allocate N MB (released after request)

cpu

CPU burn loop for N ms

leak

Allocate N MB (never freed)

error

Random 5xx WebApplicationException

exception

Unhandled RuntimeException (HTTP 500)

threadpool

Block 10 threads via CountDownLatch

contention

10 threads competing for single synchronized lock

gc

Rapid alloc/dealloc loop causing GC pressure

intermittent

Random failures at configurable rate

deadlock

Two threads deadlocked on competing locks

Evaluation Pipeline

Each test uses an LLM-as-judge approach: a scorer LLM evaluates the analysis produced by the AI provider LLM.

sequenceDiagram
participant Test as Test Case
participant App as Companion App
participant AI as AI Module
participant Scorer as Scorer LLM
Test->>App: Inject chaos
Note over Test: Wait for telemetry ingestion
Test->>AI: GET /analyze/n
AI-->>Test: Analysis + captured tool outputs
Test->>Scorer: Evaluate analysis against telemetry data
Scorer-->>Test: Score (0-100) + prompt suggestions
Note over Test: Assert score >= 70

The scorer evaluates on four dimensions (each 0-25 points): completeness, accuracy, correlation quality, and actionability. If the score falls below 70, the scorer also outputs concrete prompt improvement suggestions for the system prompt.

Cross-LLM Combinations

The analyzer LLM and scorer LLM can be different providers, preventing self-evaluation bias:

./run-integration-test.sh chaos openai grok     # AI=OpenAI, scorer=Grok
./run-integration-test.sh chaos grok openai     # AI=Grok, scorer=OpenAI
./run-all-integration-tests.sh chaos            # all 4 combinations

Environment Variables

Required for integration tests:

export OPENAI_API_KEY=sk-...       # required for openai profile
export GROK_API_KEY=xai-...        # required for grok profile

Optional scorer overrides:

export SCORER_BASE_URL=https://api.x.ai/v1
export SCORER_API_KEY=$GROK_API_KEY
export SCORER_MODEL=grok-3-mini