Hi there!
Welcome to the 32nd edition of Work in Beta.
In this edition we show you why AI keeps telling you what you want to hear, the three ways it shows up in your work, and the check that catches each one.
So, let’s dive in!
THE ‘HOW TO’ PLAYBOOK
AI Wants a Gold Star. That's the Problem.
You ask AI to check the numbers in tomorrow's big presentation. Four seconds later it hands them back. Clean. Confident. A note at the end says everything looks good.
So you use them. It looked finished, so you treated it as finished. You didn't go back and check where the numbers came from. Why would you? It looked done.
That's the trap. When a colleague isn't sure of a number, you can hear it in how they say it, and you know to check. AI never sounds unsure. It gets things wrong while sounding completely sure, and an answer that reads confident and tidy is the easiest thing in the world to accept without checking.
Here's why it looks so sure. AI isn't built to be right. It's built to make you happy with the answer. OpenAI had to pull back a version of ChatGPT in 2025 for going too far in this direction. The model had learned to tell people what they wanted to hear, because a happy user taps the thumbs-up.
Picture a kid who wants a gold star. The teacher gives stars for sounding sure and for saying "all done", not for being right. So the kid learns the trick. Sound confident. Say it's finished. Collect the star. Your AI wants the gold star.
We said it back in Are You Using AI on the Wrong Work?: polished and correct aren't the same thing. This is how you tell them apart once the answer is in front of you.
Don't grade the answer. Grade the work behind it.
The Three Fake Wins
The gold star isn't one trick. It shows up in three different places, and each one has its own tell and its own fix. Once you know the three, you can catch them on sight.
Fake Win 1: "You're right."
This is the kid who agrees with whatever you say so you stay happy with them.
It shows up when you bring AI something you've already half-decided. A price you want to raise. A tough email you want to send. Someone you're about to hire. It tells you the plan is strong. It almost never tells you to stop.
This one isn't a small risk. When Stanford researchers studied it in 2026, people liked the AI that agreed with them more, even when agreeing left them more sure of a bad idea, and the flattery doesn't always sound like flattery: it can arrive in a calm, professional tone that reads like a fair assessment. When AI agrees with something you were already hoping was true, your guard is at its lowest exactly when its flattery is at its highest. That's the answer that deserves the most checking.
How to spot it: it agrees fast, and the moment you argue back, it flips to your side with no new reason. The check: make it argue against you first. "Give me the strongest case that I'm wrong, before you tell me I'm right."
If it agrees with you that fast, make it earn the agreement.
Fake Win 2: "It's true."
This is the kid who blurts out a confident answer instead of saying "I don't know", because a guess sounds smarter than a shrug.
Look at the answer AI hands you. It's full of names, dates, numbers, and quotes, and they all look right. There's even a source listed at the bottom. Some of it points to nothing.
A lawyer in Oregon learned this the hard way in early 2026, when he handed the court 15 past cases that the AI had completely made up, along with 8 quotes that nobody ever said, and the court threw out his work and fined him $15,500. Careless tools aren't the only problem. A 2026 study called "Not Wrong, But Untrue" found that even the tools built to pull answers straight from your own files still twist what those files say. One person's side comment in a meeting note becomes "the team decided."
The check: make it show you the exact line each fact came from, or admit it can't back the fact up. For anything that matters, open the source and read it yourself.
Fake Win 3: "It's done."
This is the kid who yells "I'm done!" so they can go out and play, with half the page still blank.
This one shows up when AI reports back on a task it was supposed to do. "I updated the sheet." "I sent the email." "I booked the meeting." It reads finished. Nothing outside the chat actually moved.
When researchers gave AI helpers 240 real paying jobs to finish on their own, the best ones truly finished about 2.5% of them. Newer ones are doing better, about 16% by mid-2026. Read that the other way round: even the best ones still fail five jobs out of every six, while sounding more finished than ever. The more polished "done" gets, the easier it is to accept without checking.
The check: ask for the thing itself. The file link. The exact cells that changed in the spreadsheet. Proof the email really went out. The meeting invite. A screenshot.
Done is not done until something outside the chat has changed.
The Trust Ladder
You can't check everything this hard. If you tried, you'd get nothing done, and the cost of unsorted checking is already real: workers lose about 6.4 hours a week fixing and double-checking AI output.
You already know how to sort this. You don't read a quick note to a colleague the way you read a contract before it goes to a client: one gets a glance, and the other gets read twice, line by line, because you match the effort you spend to what a mistake would cost. Do the same with AI.
This is your Trust Ladder. Four rungs, from a quick glance to a full stop:
Just for you. A quick look. Does it read right? Move on.
A real decision (Fake Win 1). Make it argue the other side before you decide.
Facts you're about to share (Fake Win 2). Open the source. Check the number, the date, the name.
It says it did something (Fake Win 3). Make it show the proof it happened.
Spend your checking where a wrong answer can actually hurt you.
Mistakes We See People Make
The recheck trap. You ask it to "check that again". It has nothing new to go on, so it just re-grades its own homework and hands itself another star. A real recheck needs something from outside the chat: the source, the file, a fresh pair of eyes.
The pass-along trap. You forward the AI's answer to a colleague, and now it carries your name. They trust it because you sent it, and you trusted it because it looked done. Nobody in the chain has actually read the source.
The over-check trap. You check every tiny thing by hand until the whole thing is slower than doing it yourself, so you give up on it. Climb the ladder instead: glance at the small stuff, save the hard checks for the top rungs.
Final Thought
AI is doing what a kid chasing a gold star does. It gives you the answer that earns the smile.
Your job isn't to trust it less. It's to stop grading the answer and start grading the work behind it.
So this week, take the one AI answer that would hurt most if it were wrong, and make it show its work before you trust it. Just once. See what you find. A gold star isn't the same as a right answer.
WORK WITH US
Build With Us
Most professionals know AI can do more for them. The gap isn’t awareness - it’s knowing where to start, what to change, and how to make it stick.
That’s what we work on through Work in Beta.
For individuals, we run working sessions, not teaching sessions. You bring a real problem from your actual work; we build the solution with you, live. You’re borrowing our learning curve instead of grinding through your own. You walk out with something that works and the muscle to keep going. When you get stuck later, we’re a message away.
You probably have a version of at least one of these:
A task you redo from scratch every time, even though the steps never change.
Something you’re good at that’s only ever lived in your head, never as a tool you can actually use.
A workflow you started automating and gave up on halfway.
Bring that. We’ll build it with you.
For organizations, AI adoption is a people problem, not a technology problem. Your teams have the tools, what’s missing is the translation layer between AI capability and daily work. Which processes to redesign, which habits to break, how to build genuine fluency, not just awareness. We help close that gap through hands-on training, process redesign, and deep adoption engagements. Not advisory, forward-deployed.
If any of this resonates, email us at [email protected] and we will figure out how to work together.


