Framework guide

Test your LlamaIndex agent

LlamaIndex powers RAG and query pipelines over your own data, so the failure usually isn't the model, it's what got retrieved. Connect yours to ClientCoded and every query, retrieval, and synthesis step gets captured and scored, so you see when the answer is wrong, not just fluent.

1. Install

Add the ClientCoded package to your LlamaIndex project.

Terminal
$ pip install clientcoded

2. Initialize

Add two lines at the top of your LlamaIndex app, before your agent runs. Your agent_id and team api_key come from your dashboard.

Python
import clientcoded
clientcoded.init(agent_id="your-agent-id", api_key="your-team-api-key")
# Every LlamaIndex call is now captured and scored.

3. What gets auto-traced

Once init runs, ClientCoded auto-instruments LlamaIndex. You do not change your agent code. On every run, these are captured automatically:

  • Query-engine calls
  • Retrieval and vector-store lookups
  • Node post-processing and reranking
  • LLM synthesis calls
  • Agent and tool calls (for LlamaIndex agents)

4. See your results

Each run is scored and sent to your dashboard, with the overall grade, the per-dimension breakdown, and the exact step where your agent went wrong.

Output
✓ Connected. Tracing LlamaIndex calls.

Run scored: 7.8 / 10
15 steps traced · 2 issues flagged
# View the full breakdown in your dashboard

Open your dashboard to see traced runs, scores, and flagged failures. From there you can run the full adversarial test suite against your LlamaIndex agent and track quality on every change.

Next steps

See the full integration docs for the production monitoring webhook and the Query API, or start free, your first adversarial test is on us.

Testing another framework? LangChain · CrewAI · AutoGen