Dispatch / tasks
Benchmarks & comparisons
Head-to-head runs, and what the numbers actually say.
6 notes · recently updated first
- 2026-10-07 Layered Memory for Grok Build: Native Notes + Repo Files for Engineering Teams Grok Build's native memory handles cross-session recall. We layered a repo-file system on top for git-auditable team memory and ran both configurations through the same five tasks: what each layer covers, a phantom-memory boundary, and the habit that fixes it. Playbook
- 2026-10-06 We Ran Our Memory System Inside Hermes. It Already Had One. Hermes Agent grows its own memory: sessions get distilled into skills over SQLite FTS, tended by a Curator. We installed our file-based memory discipline into v0.21.5 anyway. Three verified mechanics, one caveat from the issue tracker, and a clear answer on which layer does what. Playbook
- 2026-10-06 OpenHands Model Router on your own endpoint: what works, what isn't wired yet, what it costs Day-one field test of OpenHands' Model Router on a custom endpoint: routing works at $0.0025 a call, but two early-build gaps are worth knowing first. Guide
- 2026-10-05 Where to get Kimi K3: every access route compared Kimi K3 access routes compared: official API, inference clouds, gateways, subscriptions, free tiers and self-hosting, with current prices and who each route fits. Guide
- 2026-10-01 AI Memory: Three Routes, Three Bills Three routes to AI memory — hosted API, self-hosted engine, plain files — and what each one bills you. We benchmarked the hot one, then said no. Playbook
- 2026-09-29 AI Reviews AI: We Planted 7 Bugs in a Rate Limiter and Sealed the Answer Key One agent wrote a token-bucket rate limiter with 7 planted bugs. Another agent — Kimi K3 in OpenHands Agent Canvas — reviewed it. It caught 6, spotted the rigged test, and found 2 bugs we didn't plant. Full scorecard and the $0.16 bill. Tutorial