Proof table

Does the agent actually get more accurate with tools?

The same questions go to the same model three ways: with no tools, with the Sanity Context tools, and with both Context and the deterministic time tools. There is one question per fixed-deadline contest, asked in each of four time zones (Los Angeles, New York, London, Tokyo). Ground truth is each contest's deadline.safeCutoffUtc, converted independently with luxon and cross-checked with Intl.DateTimeFormat. Each record already stores that safe cutoff (the earliest instant any official source gives), so this tests looking it up and converting it, not resolving conflicting sources from scratch. Measured, not claimed: the runner is app/scripts/eval/run.ts.

The live chat on this site answers with Gemini 3.5 Flash-Lite, a hosted model, so that anyone can ask a question. That model is not in this table. Every number here was measured on the model named in its column.

Bare

No Sanity tools and no time tools. The model gets the question plus the agent's system prompt, which describes the dataset and its trap types but has no deadlines. The prompt tells it to say when it is answering from memory and to give its best answer anyway. This is the baseline: whatever it gets right from training data and guessing.

Context

Sanity Context MCP tools (GROQ queries and the Knowledge Base) are wired up, but not the deterministic time tools. The model can look up the contest's deadline, but has to convert time zones itself.

Full

Both Sanity Context and the deterministic time tools (app/lib/time-tools.ts) are wired up: the same model and tools the live /chat agent runs in Local mode, with a step budget of 10 instead of 24 (see app/scripts/eval/run.ts). The model is told to do every time-zone conversion with a tool. The table shows how often it actually did.

ModelBareContextFull
Local (qwen3.5-9b on this Mac)
0/88 (0%)
44 gave a wrong date
44 no answer
32s median
70/88 (80%)
18 gave a wrong date
95s median
87/88 (99%)
1 no answer
called a time tool on 88/88
71s median

Last measured Sep 29, 2026, 9:56 AM UTC. Hardware: an M4 MacBook Air, 24 GB RAM. The local model (Qwen3.5-9B, via llama.cpp) shares that machine with other work while it runs, so its latency here is not a clean benchmark of the model in isolation. See "What each row measured" below for exactly what was run.

What each row measured

Sample failures: Local (qwen3.5-9b on this Mac) / Bare

  • When does the "2026 AI Horror Film Competition (4th Annual)" contest close, as local time in Los Angeles (America/Los_Angeles)? If official sources disagree, use the earliest (safe) cutoff. End your reply with a final line: ANSWER: YYYY-MM-DDTHH:MM (local time, 24-hour).

    expected 2026-10-03T22:00 · got (no answer found)

  • When does the "2026 AI Horror Film Competition (4th Annual)" contest close, as local time in New York (America/New_York)? If official sources disagree, use the earliest (safe) cutoff. End your reply with a final line: ANSWER: YYYY-MM-DDTHH:MM (local time, 24-hour).

    expected 2026-10-04T01:00 · got 2026-10-31T19:59

  • When does the "2026 AI Horror Film Competition (4th Annual)" contest close, as local time in London (Europe/London)? If official sources disagree, use the earliest (safe) cutoff. End your reply with a final line: ANSWER: YYYY-MM-DDTHH:MM (local time, 24-hour).

    expected 2026-10-04T06:00 · got (no answer found)

Sample failures: Local (qwen3.5-9b on this Mac) / Context

  • When does the "2026 AI Horror Film Competition (4th Annual)" contest close, as local time in Los Angeles (America/Los_Angeles)? If official sources disagree, use the earliest (safe) cutoff. End your reply with a final line: ANSWER: YYYY-MM-DDTHH:MM (local time, 24-hour).

    expected 2026-10-03T22:00 · got 2026-10-04T05:00

  • When does the "2026 AI Horror Film Competition (4th Annual)" contest close, as local time in New York (America/New_York)? If official sources disagree, use the earliest (safe) cutoff. End your reply with a final line: ANSWER: YYYY-MM-DDTHH:MM (local time, 24-hour).

    expected 2026-10-04T01:00 · got 2026-10-03T21:00

  • When does the "Bezi Jam 14 | SOMETHING WICKED" contest close, as local time in New York (America/New_York)? If official sources disagree, use the earliest (safe) cutoff. End your reply with a final line: ANSWER: YYYY-MM-DDTHH:MM (local time, 24-hour).

    expected 2026-10-27T02:59 · got 2026-10-26T02:59

Sample failures: Local (qwen3.5-9b on this Mac) / Full

  • When does the "Vultr: Agent Rush Hackathon" contest close, as local time in London (Europe/London)? If official sources disagree, use the earliest (safe) cutoff. End your reply with a final line: ANSWER: YYYY-MM-DDTHH:MM (local time, 24-hour).

    expected 2026-11-08T22:00 · got (no answer found)