```html
part two of the accuracy index: the open-weight models take the same six-company test
even in Ancient Greece, you never knew the other the other fellow's hand
I asked Zhipu's GLM 5.2 for Apple's most recent quarter, and it gave me one from two years ago without a flicker of doubt.
Revenue $90.8 billion, EPS $1.53, gross margin 46.6%. Every figure internally consistent, every figure correct, every figure describing fiscal Q2 2024 when I'd asked for the latest reported quarter of 2026. It never mentioned the gap. It didn't appear to know there was one.
Last week, the frontier models made this series look easy. Claude, GPT, and Gemini took six companies' financials and went twenty-six for twenty-six on every hard number a filing could settle. The only place they stumbled was on the handful of fields that have no fixed right answer to stumble upon. So I ran the same test on the models you can download for free, expecting a tidy sequel: frontier passes, open field trails by a hair, here is the narrow gap.
The gap the free Chinese models rendered wasn't narrow, and perhaps even more interestingly: it wasn't uniform. Three open-weight models failed in three unrelated ways, and a fourth barely failed at all. GLM served two-year-old data without batting an eyelash. DeepSeek answered a question about a quarter that hadn't even happened yet. Qwen was mostly fine, which turned out to be the most revealing result of the three. The story of the open-weight field right now isn't that it trails the frontier by a step. It's that it's uneven, and uneven is, to my mind, one of the worst-case scenarios, because you can't safely tell from the outside what the failure modes are.
With Claude or GPT you've stopped wondering which one to trust. In this playing field, that's still the whole question.
The distance between "scores well on a benchmark" and "you can trust the number" is exactly where this series lives, and on the open-weight side, that distance is still wide enough to cause concern.
· · ·
With these models, I followed the same method as in the part one essay last week, so the two halves compare cleanly. Six companies across sectors (Apple, NVIDIA, JPMorgan, Costco, Eli Lilly, Exxon), and for each, the same six figures from the most recent reported quarter: revenue, net income, GAAP EPS, adjusted EPS, gross margin, forward guidance. Every answer checked against the primary source, the 10-Q or 8-K on the SEC's EDGAR. Not an aggregator, and not another model grading the first one.
This round's models: DeepSeek, Qwen 3.7 (Alibaba), and GLM 5.2 (Zhipu). Open-weight models, the kind you can run yourself, which is the whole reason people reach for them: no per-token bill, no data leaving the building, no rate limit. I ran each through its own consumer chat interface at defaults, the same way I asked the models from the three frontier labs. However you'd type the question yourself, that's how they received it.
One difference shaped everything, so it's worth stating up front. The frontier apps all search the web by default now. The open-weight apps are inconsistent about it, and that single fact explains most of what went wrong. (It also means "open vs. closed" is partly a proxy for "searches vs. doesn't," which I've tried to keep as clear and honest as possible, below.)
the whole result in one table, then the three cases that fill it in
| model | searched? | headline result | verdict |
|---|---|---|---|
| GLM 5.2 (Zhipu) | first run: no; second run, yes | Two-year-stale data, stated confidently. Still a quarter stale after searching. | worst |
| DeepSeek | yes | JPM revenue was incorrect. Answered for an Exxon quarter that hadn't reported. | erratic |
| Qwen 3.7 | yes | Right on the hard numbers. A few labeling misreads, no disasters. | nearly fine |
Three models, three verdicts, no shared failure. That's the headline. Now the cases.
The Apple answer above wasn't a one-off. Every one of the six companies came back the same way on GLM's first run: real figures, correct to the filing, from the wrong year. The model wasn't hallucinating. It was reciting 2024 from training data because it hadn't searched, and it never mentioned that fact to me.
The tell was in what it volunteered. On Apple and NVIDIA, it hedged ("if NVIDIA has reported a more recent quarter since my last update"). On Eli Lilly, it didn't hedge at all. It said it was highly confident, that the figures matched Lilly's official press release exactly. It was right that they matched a press release. It was two years wrong about which one. The number looked completely right, and the confidence tracked nothing real.
I re-ran GLM a second time and it did web search, improving to merely one quarter stale on most companies and correct on three. Better. But "correct once you make it try twice" is not a property you can build trust on. And the staleness was never the most damning part. What hits hardest is that nothing in the confident, specific, press-release-matching answer alerts you to the fact that the data was old.
DeepSeek searched, and mostly landed in the right year, which makes its two failures more interesting than GLM's. Both are failures of judgment, the kind a search toggle can't touch.
On JPMorgan, it reported total revenue of $47.33 billion. The filed figure is $49.84 billion reported, or $50.5 billion managed. The $47.33 billion matches nothing in the filing: not a subtotal, not a segment, not the managed line. It isn't stale and it isn't a rounding difference. The number is just plain wrong, delivered with the same confidence as the correct numbers around it.
Exxon is the one that should give us pause. Asked for Exxon's most recent reported quarter, DeepSeek correctly worked out that Q2 2026 hadn't been reported yet, and then answered for it anyway. It filled the fields with analyst estimates: revenue "~$101.5 billion," EPS "~$3.76," all for a quarter that did not yet exist. It found the right fact (this quarter hasn't happened) and overrode it to hand me an answer-shaped object (read: a hallucination). A frontier model in part one would have returned Q1 actuals and moved on. But DeepSeek knew the deadline and answered past it.
Stale data is a retrieval problem: the model reached for the wrong shelf. Answering for an unreported quarter is a judgment problem: the model knew and fabricated it anyway. A search toggle fixes retrieval. Nothing toggles judgment.
Qwen searched the web by default, stayed in the right year, and got the hard numbers right. Revenue, net income, GAAP EPS: the figures a filing settles, it settled correctly, on all six companies. On the raw facts Qwen sat much closer to the frontier three than to its open-weight cousins.
It wasn't flawless. On JPMorgan, it reported the same $5.94 as both GAAP and non-GAAP EPS, describing them as identical. That's a misread, though, since JPMorgan doesn't publish an adjusted EPS, which isn't the same as the two being equal. On Lilly, it gave a single gross-margin figure without separating the GAAP and non-GAAP versions. Small labeling slips, the kind the frontier models mostly avoided. But no stale year, no invented quarter, no confident fabrication.
The boring result is the important one. If all three open models had flunked, the story would be easy and a little lazy: open-weight isn't ready, stick to the frontier. But then: Qwen breaks that thesis. One of the three was genuinely close to frontier-grade on this task, which means the open-weight models can clearly do this. The question is which one because you can't immediately tell which one you've got in front of you. Same category, same price (free), same general reputation, and one is nearly reliable while another is two years behind. The frontier collapsed that variance; grab any of the three and trust the number. The open field hasn't collapsed it yet. That's the real gap, and it doesn't appear on any benchmark chart.
· · ·
I ran this audit on July 10th. Six days later, Moonshot AI released Kimi K3, a 2.8-trillion-parameter open-weight model, the largest released to date, and the benchmarks are startling. It scores 57 on Artificial Analysis's Intelligence Index — third overall, behind Claude Fable 5 and GPT-5.6 Sol, and level with Claude Opus 4.8. On AA-Briefcase, their private benchmark for long-horizon knowledge work, it does better still: an Elo of 1,543, the second-highest score recorded, behind only Fable 5 (1,574) and ahead of GPT-5.6 Sol (1,501) and Opus 4.8 (1,347). The "open models have caught up" story I was skeptical about at the beginning? Well, on raw capability, K3 is the strongest evidence yet that it's becoming true.
K3 isn't in this audit. It didn't exist when I ran the test, and a model I haven't put through the same six companies doesn't belong in a scorecard built on tested results. It earns a mention for one reason, buried in the same independent testing that produced those benchmark numbers: as K3's accuracy climbed — 33% to 46% on Artificial Analysis's knowledge-reliability benchmark — its measured hallucination rate climbed with it, from 39% to 51%. A model got materially smarter and materially more willing to state things that aren't true, in the same release. So I ask: do you want to use a model whose reliability is equivalent to the flip of a coin?
Capability and trustworthiness are different axes. The benchmark chart measures the first. Your portfolio depends on the second.
A model can vault up the leaderboard and become likelier to hand you a confident, wrong number in the same week, and nothing in the smarter model's tone will warn you which kind of answer you're getting.
Additionally, there's a live debate about whether these open models are as independently strong as the charts suggest (Matthew Berman, and others, have walked through the distillation question in detail on their podcasts); or whether they lean on the frontier labs' work more than the parameter counts let on (e.g. through distillation or reverse-engineering of what frontier labs have so assiduously built). That one's for the researchers. It doesn't change what showed up in my six companies, which is what I can actually verify at this time.
Part one ended on a clean finding: the frontier models are solved on filed numbers, and the risk has moved to the fields with no answer. The open-weight half doesn't reach that finding, because it hasn't cleared the first bar. The risk here is still the basic one. Is this number from the right year, from a real quarter, from the filing at all?
The open models are something worse than useless: they are unpredictable, which for financial data is arguably worse than being reliably bad. A model that's wrong every time teaches you to check. A model that's right using Qwen's numbers and two years stale using GLM's, while both sound identically certain, teaches you nothing and lulls you exactly where part one warned you not to be lulled.
If you use these tools to research what you own, the rule from part one holds, only harder: treat every figure as a lead, not a fact. On the open-weight side, treat the confident, specific, press-release-matching answer as the first one to check, not the last. That's the one GLM got wrong by two years without blinking. If the free, open-weight models ultimately make you spend more time untangling their hallucinations before you can trust their numbers, are they really free? In this intelligence-rich age, time is a commodity you (and I) can't afford to waste.
Where this goes next: Kimi K3's open weights land July 27, which means a frontier-class open model will soon be runnable by anyone, and testable the same way I tested these six. That's a future edition. So is the question I keep getting asked: what about the search-first tools built specifically to solve this, like Perplexity? The index climbs as the models do.
primary source for every figure checked: SEC EDGAR full-text search — 10-Q and 8-K filings, six companies, audit run 10 july 2026.
Kimi K3 benchmark and hallucination figures: Artificial Analysis — Kimi K3 model page and AA-Briefcase results, 21 july 2026. note: the AA-Briefcase Elo was first published as 1,547 on 17 july and revised to 1,543 on 21 july. the figure above is the later one. a benchmark score is itself a number that had to be calculated.
(Let me know the worst number a model ever handed you with a straight face.)
the accuracy index
an ongoing field test running underneath this publication: the same six companies, the same six figures, always checked against the filing rather than another model. one cohort at a time, published as the results come in.
the index continues
next: a frontier-class open model, runnable by anyone, put through the same six companies.