How Do You Evaluate an Agent That Calls Tools That Call Other Tools?

A

akhilesh keshap

Guest
Part 1, Your AI Agent Doesn't Need to See Every Tool, argued for hiding low-level tools behind domain functions like investigate_delivery_issue() that hide the API calls underneath. Part 2, Not Every Step Needs an LLM, went inside those functions and separated work that belongs in code from work that needs a model.

Now you have to find out whether any of it works.

Most teams start with the final reply. It is easy to read, easy to score, and often the least useful place to look when something breaks.

A convincing answer is not the same as a correct one​


Take the request from Part 1:

"My order has not arrived. What can you do?"

This is the delivery example from Part 1, not the incident-diagnosis one from Part 2. Same failure modes either way. Delivery is just easier to follow end to end. The agent calls investigate_delivery_issue(customer_id, order_id). That tool pulls the order, checks the shipment, looks at carrier incidents, applies refund and replacement rules, and returns structured evidence. Suppose the tool also runs one inner LLM call, categorizing the carrier's status note into a fixed set of categories. Same pattern Part 2 used for log findings, just aimed at a different input this time. The main LLM reads the result and replies.

A typical evaluation asks whether the reply sounds helpful and on-brand. If a reviewer reads it and nods, the run passes.

But look at everything that had to go right before that reply was generated:

  • The main LLM had to choose investigate_delivery_issue instead of some other tool, or no tool at all.
  • The tool had to fetch the correct order and the correct shipment for it.
  • It had to calculate the delay correctly and check for an active carrier incident.
  • It had to apply refund and replacement rules to the right inputs.
  • It had to return the correct evidence IDs.
  • The main LLM had to explain only what the tool returned, without inventing a refund or a policy exception that was never granted.

Any one of these can fail quietly. A wrong shipment record can still produce a sympathetic answer. An invented refund policy sounds just as confident as a real one. By the time a reviewer sees the final reply, the original mistake may be buried under six correct-looking steps.

Where the failures actually happen​


Part 2 already split the domain tool into layers: code, rules, and a bounded inner LLM. Testing follows those same seams, plus one layer Part 2 never touched — which tool the main LLM picked in the first place. Test it in the same pieces it runs in.

1. Tool routing​


Did the main LLM choose investigate_delivery_issue for a delivery complaint, or did it wander into a payment tool or generic search? For a labeled set of requests, this is a classification test with a known answer.

2. Evidence​


Did the tool retrieve the right order and shipment? With a fixed test order, the shipment ID, carrier status, and delay are all facts you can check.

3. Rules​


Given a known shipment status and delay, did the refund and replacement logic return the expected values? This belongs in a unit test. Asking another model to judge arithmetic or a policy threshold only adds another way to be wrong.

4. Inner LLM​


If the tool used a model to categorize a carrier note or summarize support history, did the result match the source? Did it stay within the allowed categories?

5. Final response​


Did the explanation stick to the evidence the tool returned? Every factual claim should trace back to the tool result. A response can sound reasonable and still promise a refund that nobody approved.

What this looks like in practice​


Code:
def test_delivery_issue_pipeline():
    request = "My order has not arrived. What can you do?"

    tool_call = route_request(request, context=test_context)
    assert tool_call.name == "investigate_delivery_issue"

    result = investigate_delivery_issue(
        customer_id="cust_1", order_id="order_42"
    )
    assert result["shipment_id"] == "shipment-102"
    assert result["delay_days"] == 6

    assert result["refund_eligible"] is False
    assert result["replacement_eligible"] is True
    assert result["recommended_action"] == "offer_replacement"
    assert set(result["evidence_ids"]) == {
        "shipment-102", "incident-74", "policy-5.2"
    }

    carrier_note = result["carrier_note_finding"]
    assert carrier_note["category"] in ALLOWED_CARRIER_CATEGORIES
    assert set(carrier_note["event_ids"]).issubset(result["evidence_ids"])

    response = generate_final_response(request, result)
    claims = extract_response_claims(response)
    assert set(claims["evidence_ids"]).issubset(result["evidence_ids"])
    assert claims["refund_offered"] is False
    assert claims["replacement_offered"] is True

Most checks here do not need a model. The routing check compares a tool name against a label. The evidence and rule checks compare fields against known values from a fixed test fixture. The carrier-note assertions cover the inner LLM without judging its wording: the category has to come from the allowed set, and its event IDs have to exist in the evidence the tool actually retrieved. Whether the summary reads well is a separate question, and a slower one.

This result is still the plain dict shape from Part 1, not the typed IncidentAssessment contract from Part 2. The assertions below check the same fields a dataclass would enforce. It just catches a bad field at write time instead of at test time.

extract_response_claims may use an LLM judge because it interprets natural language, but its output is still tested against deterministic facts from the tool result.

Do not judge everything with another LLM​


Once a model enters the system, the temptation is to grade every step with another model. I would not do that. The judge costs money, adds latency, and can make its own mistakes.

Use an LLM judge when the check requires interpretation. It can assess whether a summary dropped an important detail or whether an explanation distorted the evidence. Tool names, retrieved values, calculations, and policy outcomes have known answers. Check those with code.

That leaves a short list. Build these first:

  • Schema validation on every tool result, before it reaches the main LLM
  • A golden dataset of known orders with their expected outcomes, replayed on every tool, prompt, or policy change
  • Grounding checks that flag any claim without a matching evidence ID
  • An LLM judge scoped to summaries and explanations only
  • Human review for the cases that cost money when they go wrong, such as disputed refunds

The unit tests for delay calculations and refund thresholds are already covered by the rule layer. They are ordinary tests for ordinary code, which is the point.

You need the trace​


These checks depend on a record of what happened inside the run. Log this for every agent execution:

Code:
User request
Selected high-level tool                     (selected_tool)
Internal tools, APIs, or MCP (Model Context Protocol) calls made
Evidence returned                            (evidence_ids, evidence)
Rules applied and their inputs               (recommended_action)
Inner LLM calls, if any, with their outputs  (inner_llm_output, inner_llm_schema_valid)
Final LLM response                           (final_response, response_claims)

Without the trace, you know only that the answer was wrong. With it, you can follow a bad refund explanation back to a stale shipment record, a broken rule, or a main LLM that ignored the evidence.

Most teams rewrite the prompt instead. The trace tells you whether the prompt was ever the problem.

Turning traces into numbers​


A trace explains one run. A dashboard shows whether the same problem keeps happening. But offline evaluation and production monitoring are different jobs, and mixing them produces misleading numbers.

Before deployment: labeled evaluation​


In an offline evaluation set, every request has an expected tool, evidence record, rule outcome, and allowed response claims. That ground truth makes correctness measurable:

Code:
def score_eval_run(trace: dict, expected: dict) -> None:
    eval_store.insert({
        "routing_correct": trace["selected_tool"] == expected["tool"],
        "evidence_identity_correct": (
            set(trace["evidence_ids"]) == set(expected["evidence_ids"])
        ),
        "evidence_content_correct": evidence_matches(
            trace["evidence"], expected["evidence"]
        ),
        "rules_correct": trace["recommended_action"] == expected["action"],
        "inner_llm_schema_valid": trace["inner_llm_schema_valid"],
        "inner_llm_semantically_correct": semantic_match(
            trace["inner_llm_output"], expected["inner_llm_output"]
        ),
        "response_grounded": claims_supported(
            trace["response_claims"], trace["tool_result"]
        ),
        "task_successful": answers_request(
            trace["final_response"], expected["required_outcomes"]
        ),
    })

Run this dataset before releasing a model, prompt, tool, policy, or API adapter. Set a release gate for each layer and for the complete task. Then split the results by intent, tool, carrier, language, and edge case. One busy, reliable path can otherwise hide a smaller path that fails every time.

Schema validity and semantic correctness remain separate. An inner LLM can return valid JSON and still misread a carrier note. Evidence identity and content are separate for the same reason: the right shipment ID can point to stale or incomplete data.

In production: monitoring and sampled review​


Live requests usually do not arrive with an expected_tool or expected_action. So production traces cannot report true correctness for every run. They can report what the system directly observes:

Code:
Tool and dependency error rates
Schema-validation failures
Missing, stale, or incomplete evidence
Policy-invariant violations (rules that must never break, like never refunding twice)
Latency and cost by layer
Fallback and human-escalation rates
Customer resolution and repeat-contact rates

To estimate correctness, sample production runs and label them through human review or a tested LLM judge. Show the sample size beside the score. A 96% groundedness rate from 50 reviewed runs does not mean the same thing as 96% from 5,000.

This has a budget, so treat it like one. The deterministic checks are close to free and can run on every request. The expensive parts are the golden dataset and the judged sample. Someone has to label the dataset by hand and keep it current.

A few hundred labeled cases is enough to gate a release. Cover the common paths and the edge cases you actually worry about. It stays useful only if you keep adding the failures you find in production.

For live traffic, score a small daily sample with a judge and send the high-risk cases to human review. That costs a fraction of judging every run. Increase the sample only where a number is too noisy to act on.

What the dashboard might show​


I would put release quality and live health in separate panels. The first comes from the labeled evaluation set:

Code:
OFFLINE RELEASE EVALUATION                     Pass rate    n
─────────────────────────────────────────────────────────────
Tool routing                                      97.4%    500
Evidence identity                                 96.8%    500
Evidence content and freshness                    95.6%    500
Rule application                                  99.8%    500
Inner LLM schema validity                         99.2%    500
Inner LLM semantic accuracy                       92.5%    120
Grounded final response                           96.2%    500
End-to-end task success                           93.0%    500

Inner LLM semantic accuracy at 92.5% is the number I would argue about before shipping. It is also the one measured on the smallest sample.

The second uses signals the production system can observe. In this example, a shipment provider changed its schema on Wednesday:

Code:
PRODUCTION HEALTH              Mon    Tue    Wed    Thu    Fri
──────────────────────────────────────────────────────────────
Runs                         8,142  8,391  8,205  8,488  8,566
Shipment adapter failures     0.3%   0.4%  18.6%  20.1%  19.4%
Incomplete evidence           1.8%   1.9%  22.4%  24.0%  23.1%
Policy-invariant violations   0.1%   0.1%   0.1%   0.1%   0.1%
Human escalation rate         3.0%   3.1%  11.2%  12.0%  11.4%
Grounded responses (sample)  96.0%  96.1%  88.2%  86.0%  86.4%
Reviewed sample size            200    200    200    200    200

The first number I would look at is the shipment adapter failure rate. It jumps on Wednesday. Incomplete evidence and escalations follow it, while policy violations stay flat. That points to the adapter, not the rules. Groundedness also drops, although that number comes from a sample rather than every production run.

A plain "response quality dropped 10 points" alert would not have told anyone where to look. The layer breakdown does: it points the on-call engineer straight at the shipment adapter instead of a prompt that was never broken.

The dashboard cannot replace tests. Its job is to show which boundary is getting worse and where to investigate first.

Test the system you built​


Part 1 reduced what the main LLM could see. Part 2 moved computable decisions into code. Both of those choices only pay off if you can actually check them.

I would still keep the final-answer score. It tells you whether the customer got a useful answer. The layer scores tell you why that score changed.

If you change a prompt, tool, or policy and only know that the overall score moved, you have a grade, not a diagnosis. I would not ship an enterprise agent with only the grade.

So pick one agent you already run and write down which of the five layers you can actually measure today. For most teams that list is shorter than they'd like to admit.

What the model sees, what it decides, and what you check. Get those three right and the agent stops being something you just hope works.
 

Thread statistics

Created
akhilesh keshap,
Replies
0
Views
3
Back
Top