The Authors Guild filings are an API procurement problem

This week, court documents in Authors Guild v. OpenAI were unsealed, and they are blunt. An internal OpenAI note from August 2019, cited in the plaintiffs' filings: "We trained GPT-3 on pirated stuff! No sharing that!" OpenAI policy director Jack Clark, in May 2020 testimony quoted in the same briefs, acknowledged that genre fiction authors would worry about being substituted — and that "there will be a point where a bunch of artists express worry about what we're doing here and we'll likely ignore their concerns and release anyway." The plaintiffs further allege a directive called "Project Clear" aimed at erasing evidence of data sources. Sources: Publishers Weekly, the Authors Guild, and the Hacker News thread currently on the front page.

The case will take years. That is the lawyers' timeline. If you buy API access, you are on a different one — you have to decide whose infrastructure your product sits on this quarter. So skip the legal analysis and treat the filings as what they practically are: a procurement document. Two questions to ask any model provider, in order.

1. Provenance: what was this model trained on?

For closed models, the honest answer has always been "trust us." This week put a price on that answer. The datasets at issue — Books1, Books2, LibGen-sourced corpora — were internal details customers could not inspect, and the unsealed notes suggest that was deliberate.

Open-weight models do not automatically have clean training data; openness is not a compliance certificate. But the question is askable and checkable. Weights ship with a model card and a technical report, the releasing lab puts its name on both, and anyone can probe the artifacts. When the answer to "what's in the training set?" is documented in public, nobody has to write "no sharing that!" in an internal note.

That is the entire point of procurement due diligence: not whether a vendor promises to be careful, but whether the answer can be examined without the vendor's permission.

2. Retention: what does the provider do with your data?

The filings are about training data going in. The mirror question is about your prompts going out. Two providers can serve the same benchmark scores with opposite data practices, and the difference never shows up in a latency chart.

Ask for specifics, in writing: Are prompts logged? Content, or just token counts? Are they used for training? For how long is anything kept, and can you verify that independently?

Our own answers, since we are asking you to ask: Wallaby does not log prompt content — metering records token counts, not text — and does not train on customer data. This is written into our privacy policy and is checkable behavior, not a marketing line: the billing pipeline has nothing to store the text in.

What this week changed

The revelation is not that a frontier lab cut corners — anyone following the lawsuits suspected that. It is that "knew it was illegal and did it anyway" now exists in the vendor's own words, in court documents, quotable in procurement reviews. If your compliance posture or your customers' expectations include "our AI stack is clean," an uninspected closed model just became harder to sign off.

Disclosure: Wallaby Token sells pay-as-you-go API access to kimi-k3, an open-weight model from Moonshot AI — we have a commercial interest in you asking provenance questions. We wrote this anyway, because the filings change what "trust your provider" means, and the change favors stacks you can inspect.

If you want to run the comparison yourself: sign up comes with a $0.50 trial credit, enough for a real evaluation, not a demo.