Proof table
Does the agent actually get more accurate with tools?
The same questions go to the same model three ways: with no tools, with the Sanity Context tools, and with both Context and the deterministic time tools. There is one question per fixed-deadline contest, asked in each of four time zones (Los Angeles, New York, London, Tokyo). Ground truth is each contest's deadline.safeCutoffUtc, converted independently with luxon and cross-checked with Intl.DateTimeFormat. Each record already stores that safe cutoff (the earliest instant any official source gives), so this tests looking it up and converting it, not resolving conflicting sources from scratch. Measured, not claimed: the runner is app/scripts/eval/run.ts.
The live chat on this site answers with Gemini 3.5 Flash-Lite, a hosted model, so that anyone can ask a question. That model is not in this table. Every number here was measured on the model named in its column.
Bare
No Sanity tools and no time tools. The model gets the question plus the agent's system prompt, which describes the dataset and its trap types but has no deadlines. The prompt tells it to say when it is answering from memory and to give its best answer anyway. This is the baseline: whatever it gets right from training data and guessing.
Context
Sanity Context MCP tools (GROQ queries and the Knowledge Base) are wired up, but not the deterministic time tools. The model can look up the contest's deadline, but has to convert time zones itself.
Full
Both Sanity Context and the deterministic time tools (app/lib/time-tools.ts) are wired up: the same model and tools the live /chat agent runs in Local mode, with a step budget of 10 instead of 24 (see app/scripts/eval/run.ts). The model is told to do every time-zone conversion with a tool. The table shows how often it actually did.
| Model | Bare | Context | Full |
|---|---|---|---|
| Local (qwen3.5-9b on this Mac) | 0/88 (0%) 44 gave a wrong date 44 no answer 32s median | 70/88 (80%) 18 gave a wrong date 95s median | 87/88 (99%) 1 no answer called a time tool on 88/88 71s median |
Last measured Sep 29, 2026, 9:56 AM UTC. Hardware: an M4 MacBook Air, 24 GB RAM. The local model (Qwen3.5-9B, via llama.cpp) shares that machine with other work while it runs, so its latency here is not a clean benchmark of the model in isolation. See "What each row measured" below for exactly what was run.
What each row measured
Local (qwen3.5-9b on this Mac) / Bare
88 questions scored under the "bare" configuration, drawn from the public Sanity dataset's fixed-deadline contests x 4 time zones (America/Los_Angeles, America/New_York, Europe/London, Asia/Tokyo). Ground truth is each contest's deadline.safeCutoffUtc converted with luxon and truncated to the minute, cross-checked independently with Intl.DateTimeFormat (see app/scripts/eval/questions.test.ts). Scoring is exact to the minute: an answer that rounds a :59:59 cutoff up to the next minute is later than the cutoff, so it counts as wrong.
Local (qwen3.5-9b on this Mac) / Context
88 questions scored under the "context" configuration, drawn from the public Sanity dataset's fixed-deadline contests x 4 time zones (America/Los_Angeles, America/New_York, Europe/London, Asia/Tokyo). Ground truth is each contest's deadline.safeCutoffUtc converted with luxon and truncated to the minute, cross-checked independently with Intl.DateTimeFormat (see app/scripts/eval/questions.test.ts). Scoring is exact to the minute: an answer that rounds a :59:59 cutoff up to the next minute is later than the cutoff, so it counts as wrong.
Local (qwen3.5-9b on this Mac) / Full
88 questions scored under the "full" configuration, drawn from the public Sanity dataset's fixed-deadline contests x 4 time zones (America/Los_Angeles, America/New_York, Europe/London, Asia/Tokyo). Ground truth is each contest's deadline.safeCutoffUtc converted with luxon and truncated to the minute, cross-checked independently with Intl.DateTimeFormat (see app/scripts/eval/questions.test.ts). Scoring is exact to the minute: an answer that rounds a :59:59 cutoff up to the next minute is later than the cutoff, so it counts as wrong.
Sample failures: Local (qwen3.5-9b on this Mac) / Bare
When does the "2026 AI Horror Film Competition (4th Annual)" contest close, as local time in Los Angeles (America/Los_Angeles)? If official sources disagree, use the earliest (safe) cutoff. End your reply with a final line: ANSWER: YYYY-MM-DDTHH:MM (local time, 24-hour).
expected 2026-10-03T22:00 · got (no answer found)
When does the "2026 AI Horror Film Competition (4th Annual)" contest close, as local time in New York (America/New_York)? If official sources disagree, use the earliest (safe) cutoff. End your reply with a final line: ANSWER: YYYY-MM-DDTHH:MM (local time, 24-hour).
expected 2026-10-04T01:00 · got 2026-10-31T19:59
When does the "2026 AI Horror Film Competition (4th Annual)" contest close, as local time in London (Europe/London)? If official sources disagree, use the earliest (safe) cutoff. End your reply with a final line: ANSWER: YYYY-MM-DDTHH:MM (local time, 24-hour).
expected 2026-10-04T06:00 · got (no answer found)
Sample failures: Local (qwen3.5-9b on this Mac) / Context
When does the "2026 AI Horror Film Competition (4th Annual)" contest close, as local time in Los Angeles (America/Los_Angeles)? If official sources disagree, use the earliest (safe) cutoff. End your reply with a final line: ANSWER: YYYY-MM-DDTHH:MM (local time, 24-hour).
expected 2026-10-03T22:00 · got 2026-10-04T05:00
When does the "2026 AI Horror Film Competition (4th Annual)" contest close, as local time in New York (America/New_York)? If official sources disagree, use the earliest (safe) cutoff. End your reply with a final line: ANSWER: YYYY-MM-DDTHH:MM (local time, 24-hour).
expected 2026-10-04T01:00 · got 2026-10-03T21:00
When does the "Bezi Jam 14 | SOMETHING WICKED" contest close, as local time in New York (America/New_York)? If official sources disagree, use the earliest (safe) cutoff. End your reply with a final line: ANSWER: YYYY-MM-DDTHH:MM (local time, 24-hour).
expected 2026-10-27T02:59 · got 2026-10-26T02:59
Sample failures: Local (qwen3.5-9b on this Mac) / Full
When does the "Vultr: Agent Rush Hackathon" contest close, as local time in London (Europe/London)? If official sources disagree, use the earliest (safe) cutoff. End your reply with a final line: ANSWER: YYYY-MM-DDTHH:MM (local time, 24-hour).
expected 2026-11-08T22:00 · got (no answer found)