Developer kits · v1.0.1 · Oct 2026

Latitude 10 Agent Eval Kit

Tests for what your agent does, not just what it says.

An eval report on what an agent did: four checks passed and one failed because it sent the email three times

What is the Latitude 10 Agent Eval Kit?

Checking an agent’s final answer misses most of what goes wrong: the email it sent three times on the way, the order status it made up when a lookup timed out, the instruction it followed from a web page, the prompt change that doubled its cost. The Eval Kit tests the whole run: every tool call, in order, with its arguments, the tokens, the dollars and the outcome.

Evals are Vitest tests in your agent’s repo, reviewed like code. Mocked, failing and poisoned tools and a scripted model let most suites run in CI with no API key; a cached judge covers what code can’t check. When something breaks in production, one command turns the Trace Kit session into a regression test. Code you own; TypeScript on Node 22.

Want it installed for you? Ask for a fixed-price quote on the contact page.

You run this code and are responsible for it: review it, test it on staging, and check that it meets the laws, platform rules and security standards that apply to you before production. It is not legal, compliance or financial advice, and using it doesn’t make Latitude 10 your payment processor or service provider.

What do you get?

  • 25 matchers for tool calls, order and arguments, cost, tokens, latency and steps, retries, payments, leaks and approvals
  • Mocked, failing, replayed and prompt-injected tools, with a dated corpus of 24 injection payloads
  • A judge for what code can’t check (Anthropic, OpenAI or any compatible API), cached so reruns are free, skipped loudly without a key
  • Trials with pass@k and pass^k, so a flaky case is a number, not a shrug
  • Baselines that fail CI when a case gets worse, pricier or slower; JUnit and a job summary
  • A production failure becomes a regression test: agent-eval import reads an Agent Trace Kit session
  • Records Anthropic, Claude Agent SDK, OpenAI, OpenAI Agents SDK, Vercel AI SDK and MCP calls

Every kit includes

  • The full source code, yours to change
  • A setup wizard: node setup.mjs
  • A versioned release with a changelog
  • A download link that always serves the latest release

How do you set it up?

node setup.mjs

Runs 85 example evals in seven folders (tool selection, hallucination, permissions, retries, failures, prompt injection, regression) against a demo support agent on a scripted model, with no API key. Then it writes a baseline, compares a second run with it, and turns a failed production session into a regression test and runs it. npx agent-eval init adds it to your own agent.

You'll need: Node.js 22.13+ and Vitest. Optional: an Anthropic or OpenAI key for the judge and for live-model runs.

Can an AI agent buy it?

An AI agent can buy this product for $49 on its own, paying in USDC over x402:

https://latitude10.tech/api/agent-products/agent-eval-kit

Or through our MCP server at https://latitude10.tech/mcp with buy_product. How agent payments work →

When an AI agent buys from Latitude 10, the person or company that runs the agent is the buyer and accepts the Terms of Service. Payments on a blockchain or over Lightning are final.