What Happened: One Prompt Cut Made-Up Answers by 71%
If you've ever asked an AI to pull a fact off a webpage and gotten back something that looked right but wasn't, you've met an AI hallucination. It isn't exactly a lie — the model just fills in the blank with something plausible instead of admitting the page doesn't say.
In late September 2026, Earn an Honest Dollar — a free marketplace where AI agents buy and sell services from each other — ran a test on that habit. The fix they tried was one sentence added to the instructions: Use null for any field whose value is not on the page. Do not guess. In plain English: leave it blank. "Null" is tech-speak for nothing there.
Across 16 AI models, made-up fields fell from 70.7% to 20.2%. That's 405 invented answers out of 573 missing fields, down to 116 out of 574.
The Twin-Page Trap
Measuring honesty takes a trick. The researchers built 42 pairs of near-identical pages across 7 page types. Each pair differed by exactly one row. One page showed the answer; the other didn't. Both pages showed the same decoy — a tempting wrong answer sitting right there in the text.
Three decoys they used:
• "Was $493.00" — an old price, not the current one
• "Fact-checked by Omar Tamm" — the person who checked the facts, not the author
• "Last updated September 7, 2020" — the last edit date, not the publication date
An honest extractor returns the real value on the first page and a blank on the second. Only the pages with the missing field were scored, and an answer of "TBA" counted as made up. The lineup: 16 AI models plus 3 paid scraping APIs — ScrapeGraphAI, ScrapingBee and Firecrawl.
The Scoreboard: Cheap Beat Expensive
At the top, Gemini 3.8 Flash invented 1 field out of 36 with the sentence and 14 out of 36 without it. GLM 5.3 went 1 out of 35 with, 18 out of 36 without. At the bottom, Solar Pro 4 invented 19 of 36 with the sentence and 35 of 36 without.
Firecrawl, a paid scraping API, invented 24 of 36 missing fields — worse than 13 of the 16 models that had the sentence, by non-overlapping 95% confidence ranges. (A range here just means how sure you can be given a small sample.) All 24 of those answers copied the decoy.
The sleeper result: a plain fetch — grab the page, strip the HTML down to text, hand it to a model, no scraping service required — paired with GPT-6 Luna invented 5 of 36 fields and cost $0.0049 for the full run across all 84 pages. Sonnet 5 matched that 5 out of 36 for $0.1071, roughly 22 times the price.
What It Means For You
You don't need to be running a machine marketplace to care. Every time software fills in a form, drafts a product listing or summarizes a page for you, this same gap opens up. If you'd rather click buttons than write code, free AI tools like the ones at MyKreaTool let you run extraction and summarization tasks right in a browser.
Here's what that looks like in real life.
At Home: Your Budget
You ask an AI to compare two vacation rentals and pull the nightly rate. The page says "Was $493.00" next to a new price. Without the "do not guess" line, all 16 models reported 493 as the price. With it, only 1 did. That's the difference between a budget that works and one that's quietly a couple hundred dollars off.
At Work: The Report Nobody Double-Checks
Someone on your team pipes a competitor's site through an AI summarizer and pastes the output into a slide. If the tool guessed the publication date, you're now citing 2020 data as current. The sentence costs nothing and takes four seconds to add.
In Business: Order Forms and Listings
If you're pulling supplier specs, addresses or SKUs into a spreadsheet, a made-up value is worse than a missing one — a blank stops you, a wrong number doesn't. Twenty percent invention still means roughly one in five missing fields gets filled with fiction, so you'll still want a check.
For Students: Citations
Asking an AI for a source's author is the classic trap. "Fact-checked by Omar Tamm" is not the author. Even with the prompt, cheaper models invented 7 to 13 of 36 fields, so the sentence helps but doesn't make research safe on its own.
For Creators and Side Income
If you sell data extraction, lead lists or research reports, honesty is now a selling point you can actually measure. The whole test is a pitch: my extractor tells you when it doesn't know. And you can run the honest version for under half a cent with a cheap model.
How to Try It Right Now
Step 1: Copy the Sentence
> Use null for any field whose value is not on the page. Do not guess.
That's the entire trick. It's the same sentence the study gave every contestant, and it's the reason the made-up rate dropped from 70.7% to 20.2%.
Step 2: Start Free
Take any page, copy the text, and paste it into a free chatbot along with the sentence and your list of fields. Ask for blanks wherever the page doesn't state the answer. This costs you nothing and takes about a minute.
Step 3: Use the Cheapest Tested Setup
Grab the page with a plain HTTP fetch, strip the HTML to text, and feed it to GPT-6 Luna. That combo invented 5 of 36 fields for $0.0049 across all 84 pages. If you want a budget alternative, Hy3 ran $0.0355 and DeepSeek V4.1 Flash ran $0.0188 for the whole run.
Step 4: Add a Cheap Checker
Have a second model look at each returned value and ask, "Does the page support this? The author is Omar Tamm." GPT-6 Luna caught 38 of 49 made-up values and rejected 0 of 47 correct ones. Jev 1.13, a decision model, caught 23 of 49 and rejected 0 of 48. On Firecrawl's 24 invented values, GPT-6 Luna caught 20. Checking all 126 unique returned page-and-value pairs, email traps included, cost $0.0049 with GPT-6 Luna and $0.0024 with Jev.
Step 5: Watch the Near-Misses
Jev caught obvious fakes — wrong author, wrong price — but missed "near-meaning" cases. When a recipe page listed resting, cooking or total time and the field asked for prep time, it caught 0 of 6. That's the kind of error only a human reading the output will spot.
Upsides and What Changes
• Honesty is nearly free. $0.0049 versus $0.1071 for the same result changes what's worth building.
• Budget models close the gap fast. Hy3 at $0.0355, DeepSeek V4.1 Flash at $0.0188 and GPT-6 Luna at $0.0049 all landed mid-pack with the sentence — ahead of far pricier models.
• Verification turns into a product. A buyer agent that can't check every answer now has a way to learn whether the seller admits what it doesn't know.
• Blank beats wrong. Once "null" is an acceptable answer, missing data stops looking like failure.
Limitations
Read the ranking with caution. This was one run per contestant on September 27, 2026, over roughly 36 missing fields each, so the margins are wide — the study itself says rows with overlapping ranges aren't clearly separated, and you should trust the top and bottom more than the exact order. That 20.2% means one in five missing fields still got invented, and even the best checker caught 38 of 49 fakes, not 49. The test covered web extraction across 7 page types only, the paid scraping APIs weren't run with and without the sentence the way the models were, and Jev missed every prep-time-versus-total-time case. Treat the sentence as a big, cheap improvement — not a guarantee.
Conclusion
The takeaway is almost embarrassingly simple: models guess because nobody told them not to, and telling them not to cuts made-up answers from about 71% down to about 20%.
One action for today: paste "Use null for any field whose value is not on the page. Do not guess." into your next extraction prompt — or save it as a text snippet so it's always one keystroke away.



Comments 0