Mohammed Mutahar
← Back to BlogLLM/Agent Evaluation

LLM/Agent Evaluation

August 20, 20264 min read

LLM/Agent Evaluation

LLMs are static, have fixed knowledge and are limited to text-to-text interactions. LLM based agents address those gaps, using LLMs as the backbone, integrating them into workflows and equipping them with tools.

Evaluating these agents isn't the same as evaluating their text output.

You need to evaluate their decision making, their action taking ability, their tool use.



LLM-Based Agents are divided into 4 main categories based on their abilities to do the following:

  • Plan
  • Use tools
  • Self-Reflect
  • Manage Memory


An LLM agent is a combination of two things: Backbone LLM + Agent Harness

  • Agent Harness
    • Runtime environment
    • Loops
    • Memory systems
    • Tool registries
    • Safeguard-rails


★ A survey on Evaluation of LLM Based agents (Yehudai et al.) is a comprehensive list of a collection of all evaluation techniques for all sorts of agents.

Take a look at this paper's Fig 1. for exact papers that talk about exact evaluation techniques.

Since I have created a multi-agent workflow for planning and research, let us look at evaluation metrics for that orchestration.

A multi-step planner is evaluated on its long-term planning capabilities (→ Plan Bench), its capability to follow structured workflows (Xiao et al.), or to manage real world planning tasks expressed in natural language (Zheng et al.).

Conductor (my multi-agent project) can be evaluated across 3 different layers:

(i) Final report: Evaluate the agent on its final output using a different LLM as a judge.

ProsCons
FastLimited insight into agent behavior
InexpensiveCannot judge intermediate decisions
Easy to integrateCannot identify failure cause
Suited for large scale monitoring and testingCannot gauge execution efficiency

(ii) Stepwise evaluation: Score tool call, info retrieval, planning step independently. Allows error localization.

(iii) Trajectory based evaluation: To evaluate how close the agent sticks to the trajectory of the optimal path/solution.

2 ways to do this:

Reference BasedReference Free
Compare the trajectory against an optimal path (exact, partial, unordered, subset matching)Use LLM-based judge to decide upon the coherence, efficiency, and goal directedness
They provide precision and reproducibility but depend on predefined expected behaviorProvide flexibility at the cost of reliability

Here is how each of Conductor's 4 agents can be evaluated individually:

  • Planner: Evaluate this agent based on how well it defines the subtasks, and how it sequences them. PlanBench is a relevant benchmark. Also verify Galileo's goal-oriented action advancement: whether the planner creates subtasks that each advance the current state towards the goal. This prevents the planner from creating a plan that looks reasonable but does not move the goal forward.

  • Researcher: Evaluate this agent on its intent recognition and tool-use (web-calls) (→ Online-Mind2Web).

  • Synthesizer: Evaluate how well it integrates findings across sources, and if its claims are faithful to the evidence it receives. Problems arise in long-range consistency and handling dynamic memory, therefore that is what you should be wary of.

  • Writer: Can use an LLM as a judge to evaluate its coherence, completeness and correctness.

There are some end to end evaluation guidelines and benchmarks too:

Deep Research Bench, Mind2Web 2, DeepScholar Bench.



★ When performing analysis on performance, decouple the LLM backbone from the agent harness. (Eg: if the writer performs better, was it the model that was improved, or the prompt?)

★ Most evaluations currently prioritize performance over speed/cost.



Now that you have an idea of how evaluation of different agents is done, here are all agents that are available for testing. Their testing techniques are explained and consolidated in Figure 1 of the paper: A Survey on Evaluation of LLM Based agents. (Yehudai et al.)

Agent Evaluation:

  • Based on Capabilities
    • Planning, Multi-step Reasoning
    • Function calling and tool use
    • Self-Reflection
    • Memory
  • Based on Agent Types
    • Web Agents
    • SWE Agents
    • Scientific Agents
    • Conversational Agents
  • Generalist Agent's Evaluation
  • Frameworks for agent evaluation
    • Development Framework
    • Gym-like Environments