Field notes on frontier open models.
Field notes on frontier open models — prices, benchmarks, and what we run in production.
- 2026-09-09 Why we built Wallaby — and where we're going Wallaby exists because access to frontier open-weight models is harder than it should be. What we run, what we charge, and reliability you can verify. PinnedCompany
- 2026-10-04 Five stages, zero replays: a trending paper just described a failure pipeline much like ours SkillRefiner refines agent skills offline from historical traces. We run a similar loop in production, built independently: five stages, real cases, and what it taught us back. Playbook
- 2026-10-04 Open-weight LLM API pricing, October 2026: Kimi K3, DeepSeek V4, GLM-5.3, Qwen3.8-Max Official per-1M-token prices for the four major open-weight model APIs, fetched from the official pricing pages on Oct 4, 2026 — plus the peak-window, context-tier and cache traps, and what a coding-agent workload actually costs on each. Guide
- 2026-10-01 AI Memory Pricing: Supermemory vs Mem0 vs Zep vs Letta AI memory pricing compared from all four official pages — Supermemory, Mem0, Zep, Letta — plus the self-hosted route we benchmarked and the free file option. Playbook
- 2026-10-01 AI Memory: Three Routes, Three Bills Three routes to AI memory — hosted API, self-hosted engine, plain files — and what each one bills you. We benchmarked the hot one, then said no. Playbook
- 2026-10-01 Your Agent Needs Sleep, Not a Bigger Context A stale card on our project board exposed the real bug: we had built our AI a perfect memory, and perfect memory was the problem. Human memory works because it forgets — so we rebuilt ours around sleep, pruning, and myelin. Playbook
- 2026-10-01 Memory Consolidation for AI Agents: A Three-Piece Playbook A working recipe for agent memory that gets more accurate over time: a replay job that decides what survives the day, a pruning pass that deletes on purpose, and a reflex layer that turns repeated procedures into muscle. Playbook
- 2026-10-01 Memory Is Not Written. It's Audited. Write discipline always loosens. Three weeks in, the log claims things are done that the state file still shows in flight — and nobody notices. The fix is not stricter discipline; it's a weekly mechanical diff between what your log claims and what your state shows. Playbook
- 2026-09-30 Microsoft put 1,024 AI agents on one task. What did it cost? Microsoft's Agensh paper scales one task from 1 agent to 1,024 with no orchestrator. The gains are real but sharply sublinear — and the token bill appears nowhere in the paper. Industry
- 2026-09-30 Why did reasoning_effort silently stop working? For 13 days we thought the model ignored reasoning_effort. The truth: our own router was dropping the parameter. A postmortem with logs, retests, and the fix. Playbook
- 2026-09-29 Your API key is probably already on GitHub Searching GitHub for OPENAI_API_KEY returns thousands of committed .env files — many still live. Five boring fixes that actually protect your keys. Guide
- 2026-09-29 It Remembers Because You Close Memory systems don't fail at install — they fail at the end of each session, one skipped write-back at a time. The fix is a two-minute closeout ritual, and it now ships as a paste-in prompt. Playbook
- 2026-09-29 AI Reviews AI: We Planted 7 Bugs in a Rate Limiter and Sealed the Answer Key One agent wrote a token-bucket rate limiter with 7 planted bugs. Another agent — Kimi K3 in OpenHands Agent Canvas — reviewed it. It caught 6, spotted the rigged test, and found 2 bugs we didn't plant. Full scorecard and the $0.16 bill. Tutorial
- 2026-09-28 OpenHands Agent Canvas: Scheduled Slack Digests Wire an Agent Canvas cron automation to Slack via an incoming webhook: a self-contained prompt, a real verified run, metered cost, and two environment notes. Tutorial
- 2026-09-28 The Authors Guild filings are an API procurement problem Unsealed briefs in Authors Guild v. OpenAI quote executives admitting they trained on pirated books. Two questions every API buyer should now ask their provider — provenance and retention. Industry
- 2026-09-25 OpenHands 'Failed to load provider connections': the OH_SECRET_KEY cause and fix OpenHands 'Failed to load provider connections'? Two agent-servers sharing ~/.openhands use different OH_SECRET_KEY values, so saved keys read empty. Fix: one value per machine. Playbook
- 2026-09-25 How to Use Kimi K3 in pi (One JSON Block) How to use Kimi K3 in pi: one JSON block in ~/.pi/agent/models.json, one exported key, done. Verified end-to-end on pi 0.86.1 with billing logs to match. Tutorial
- 2026-09-25 Your AGENTS.md Is Rotting (Here's the Fix) AGENTS.md rots by default: stale decisions, invented paths, confident lies. Structure won't fix it — a maintenance layer will. Here's ours, in production. Playbook
- 2026-09-25 'PromptTokensDetailsWrapper' object has no attribute 'cache_creation_tokens' — the OpenHands fix OpenHands: 'PromptTokensDetailsWrapper' object has no attribute 'cache_creation_tokens' — an SDK/LiteLLM version-skew bug, not your endpoint. Cause and fix inside. Tutorial
- 2026-09-23 OpenHands stuck at 'Running task'? Why the first conversation stalls — and the fix OpenHands stuck at 'Running task' / 'Waiting for task'? The call succeeded — the UI event channel is rate-limited. Fix: stop, start a fresh conversation. Why Metrics shows $0.0000, too. Tutorial
- 2026-09-21 AI remembers, you find: fixing AI memory rot AI memory rots mid-project: settled decisions get relitigated, files get invented. Wallaby-agent-rules v2 moves memory to plain, dated, sourced files. Playbook
- 2026-09-17 Kimi K3 in OpenCode: zero-config, in the catalog Run Kimi K3 in OpenCode with zero config — Wallaby is in the models.dev catalog. One env var, one model refresh, K3 is selectable. Verified end to end. Tutorial
- 2026-09-17 One endpoint, one bill: Kimi K3 for dev teams One endpoint, one bill: give every developer their own API key with a hard budget cap, all from one shared prepaid balance — no seat licenses, no surprise. Tutorial
- 2026-09-17 Union Alpha Free in OpenCode: how to try it Union Alpha Free in OpenCode: zero-cost model, 262k context, no key required. How to enable it, what we verified, and honest caveats before real work. Tutorial
- 2026-09-14 How to compare LLM API costs: measure cost per accepted output LLM API cost comparison goes past price per million tokens: retries and human fixes are the bigger bill. Measure cost per accepted output in one afternoon. Playbook
- 2026-09-13 Kimi K3 pricing compared: official list vs 17 providers Kimi K3 pricing, straight: $3.00/M input, $0.30/M cached, $15.00/M output at official list. What 17 API providers actually charge, and what one task costs. Guide
- 2026-09-13 AGENTS.md as a token budget: the forbidden list AGENTS.md is a Linux-Foundation standard read by 30+ coding agents. Used well it's a token budget — our full production playbook, forbidden list included. Playbook
- 2026-09-13 Kimi K3 in Codex CLI: profile-file setup Run Kimi K3 in OpenAI Codex CLI through Wallaby's API, using the profile file — one config file, verified end to end with a real coding task. Tutorial
- 2026-09-12 Kimi K3 in VS Code Copilot via BYOK: setup Point VS Code's Bring Your Own Key endpoint at an OpenAI-compatible API and run kimi-k3 in Copilot Chat agent mode — full setup with working configuration. Tutorial
- 2026-09-12 Kimi K3 in Cline: 5-minute BYOK setup Configure Cline to use the kimi-k3 reasoning model through Wallaby Token's OpenAI-compatible API. Working setup in under five minutes, verified end to end. Tutorial
- 2026-09-10 Kimi K3 in Hermes Agent: 4 lines of YAML Point Hermes Agent at the kimi-k3 reasoning model through Wallaby Token's OpenAI-compatible API. Verified end-to-end — chat and tool use both work. Tutorial
- 2026-09-09 Top up Wallaby Token: card, Apple & Google Pay How to add credit to your Wallaby Token account with a card, Apple Pay, or Google Pay — plus the pitfalls we hit ourselves when testing. Guide
- 2026-09-09 Kimi K3 in Claude Code: 5-minute setup Configure Claude Code to run Moonshot AI's kimi-k3 model through Wallaby Token's Anthropic-compatible endpoint, verified end-to-end. Tutorial
- 2026-09-09 Kimi K3 in Cursor: what blocks it, what works Running Kimi K3 in Cursor via an OpenAI-compatible endpoint hits two hard blocks today — here is exactly what we hit, and the setup that works instead. Tutorial
/ page:
No posts match. Try a different keyword or tag.