Our team built this blog, and that became our blind spot: when you write and maintain your own site, you stop reading it with the eyes of a first-time visitor. Internal reviews kept missing small but meaningful flaws. So we brought in an unbiased reviewer with no stake in our stack: Claude Opus 5.5, a model we don't make.
We ran the audit through our own inference gateway, so every token consumed was metered exactly like production customer traffic. It browsed our live blog via Cline, covering the homepage, topic listings, and full article pages. Total audit cost: $2.50, itemised.
The auditor surfaced six-to-eight functional site issues. Within 24 hours we shipped fixes for four findings and formally rejected two suggestions, with documented rationale for each. This post walks through the itemised cost, what we changed, and where the approach has hard limits.
Disclosure: Wallaby Token operates an OpenAI-compatible inference API. We sell Kimi K3; Claude is not part of our public catalog — we brought our own access for this experiment, and the metering pattern is the point. The goal was outsider perspective, and running the whole workflow under our own billing gives us real-world cost visibility.
What it found (the ones that mattered)
A feature promise we could not fulfill. One article closed with the prompt "leave a comment below if your error isn't listed." But we have no comment section on the blog. The auditor flagged this as our most obvious flaw: promising interactivity that does not exist. This stung, because the broken call-to-action lived on pages where we ask readers to trust our troubleshooting guidance.
Misaligned filter and category labels. Our "Set up" bucket mixed release notes, billing guides, and troubleshooting content. "Benchmarks & comparisons" contained field tests and playbooks. The auditor's observation was sharp: categories had drifted from consistent product-oriented semantics toward loose "this roughly fits" grouping.
Series content missing a top-level table of contents. Articles displayed "Part 6 of 9" with previous-/next-article navigation, but offered no way to view the full series at a glance. With two long-running series published, this created a genuine navigation gap for readers.
Verification metadata buried inside prose. Our articles included notes such as "Verified on Cline 4.1.22 · Oct 7" as ordinary paragraph text. There was no dedicated component, and "updated" and "last verified" timestamps were intermixed, making "latest" status ambiguous.
What we shipped within 24 hours
- Replaced the non-functional comment invitation with a static content-problem report link: preserving the intent for reader feedback without promising a feature we don't operate.
- Added a view-all series overview entry for multi-part article sequences.
- Rolled out a per-article "Did this help?" vote as the starting point for feedback on fast-changing technical content.
- Re-tagged eight posts to realign task-oriented filters with consistent semantics.
- The following morning, converted inline verification notes into a dedicated UI component with a structured
modelfield. We backfilled metadata across 32 existing articles, with one hard rule: we never fabricate verification metadata for posts that were never formally validated.
All these changes are live on the blog today.
What we rejected
Native public comment threads. Notably, the auditor itself cautioned against them: open developer-blog comment sections tend to accumulate long-term maintenance debt filled with "this no longer works on Windows" type reports. For our technical blog, structured lightweight signals — "did-this-help" votes and report-outdated links — are preferable to open discussion threads. We agreed and declined this change.
Config generators and interactive pricing calculators. These were well-scoped, reasonable ideas, but they constitute core product work rather than blog-site fixes. We moved them onto our product roadmap instead of rushing them in to make the audit feel fully resolved.
The honest limits
This represents one model, one single-pass review, one evening of runtime. Its subjective site ratings ("70–80 as a content site, 50–60 as a developer knowledge product") are qualitative opinion, not rigorous measurement. We treated the scores as contextual color and prioritized concrete, reproducible findings.
This single-pass audit demonstrates a workflow, not a formal benchmark measuring how well this model performs as a site auditor.
An external reviewer can only report problems visible to what it can browse; our internal backlog (pricing-cluster improvements, unfinished article series) still carries higher net business value than anything surfaced in this exercise.
One part of this experiment we would absolutely repeat is the metered-gateway routing pattern. By running the audit through our existing gateway, costs show up as real line-items: model identifier, token counts, final price per invocation. If you have never run an AI review of your own product while watching your usage meter tick, it is one of the lowest-cost forms of outside consultation you can run. The itemised receipt itself is part of the value.
Set up the same loop: any OpenAI-compatible key, your tool of choice, an itemised receipt on the first call. $0.50 free credit; top-ups start at $20