We ran an experiment last week that should worry anyone using AI-generated practice questions to study for a licensing exam.
One of our founders is a roofing contractor. He has never taken a real estate course, never read a real estate textbook, never worked a deal. We handed him a set of real estate practice questions — cold, no book, no prep — and told him to answer on gut instinct.
He scored 67%. Most state real estate exams pass at 70–75%.
He didn't secretly know real estate law. The questions leaked their own answers. And once you learn to see how, you'll spot the same leak in most of the practice material floating around — especially the kind you get by asking a chatbot to "make me 20 practice questions."
How a question gives away its answer
Look at this made-up example. It's the shape of thousands of real practice questions we've analyzed:
A contractor discovers a code violation partway through a job. What should the contractor do?
A) Ignore it and finish the job
B) Hide it from the inspector
C) Stop work, document the violation, notify the owner, and correct it per code
D) Ask the homeowner to sign a waiver
You got it right without knowing a single code section. Everyone does. The correct answer is the longest, most careful, most professional-sounding option, and the wrong answers disqualify themselves — nobody's licensing board recommends hiding things from the inspector.
Now compare:
Under the code, a temporary erosion-control measure must remain in place until the disturbed area reaches what percentage of established vegetation?
A) 50%
B) 70%
C) 80%
D) 90%
If you haven't studied, you're guessing at 25%. If you have, you know it. That's the entire difference between a question that measures knowledge and a question that measures test-taking instinct.
The giveaway patterns are consistent: the most conscientious-sounding option wins; the correct answer is the longest and most hedged; wrong answers use absolute words like "always" and "never"; an option repeats key words from the question; scenario questions reward "be responsible" instead of a fact. A question with those tells can be answered by anyone — which means it teaches you nothing and, worse, tells you you're ready when you aren't.
We measured it — including the stuff you can buy
We built a blind test: an evaluator answers each question with no access to the study material, and has to explain how it got the answer. If it can crack the question through wording tricks, common sense, or process of elimination, the question is guessable. If the only path is knowing a specific figure, deadline, threshold, or procedure — that's a real question.
Then we scored everything we could get our hands on:
- A major test provider's own prep book: about 10% guessable. That's what professional question-writing looks like — the people who write the actual exams kill leaky questions on purpose.
- Another commercial prep book, sold for real money: 58% guessable. More than half its questions can be answered without opening a book.
- AI-generated questions with no quality controls: 60%+ guessable. And we're not throwing stones — that number came from our own early question generator. When you casually ask any AI for practice questions, it reaches for the same mold every time: a responsible-sounding scenario with one obviously-professional answer. It looks like studying. It isn't.
That last point is the one to remember before you paste "quiz me for the plumbing exam" into a chatbot. The questions will come back clean, confident, and professional-looking — and mostly answerable by your instincts instead of your knowledge. You'll score 85–90%, feel ready, and walk into a PSI or Pearson VUE exam written by psychometricians whose entire job is making sure instinct isn't enough.
The exam fee, the weeks waiting for a retake slot, the jobs you can't legally take in the meantime — that's the real price of practicing on soft questions.
How to pressure-test any practice source
Five checks you can run on any prep material in about two minutes:
- Cover the book test. Could a smart friend in a different trade answer this question? If yes, it's not testing knowledge.
- Length check. Is the correct answer consistently the longest, most qualified option? Real distractors are the same length and specificity as the answer.
- Absolute words. Do wrong options say "always," "never," "only"? Test-writers' oldest tell.
- Numbers and deadlines. Good questions pin specific figures — days to file, minimum coverage, net worth requirements, code thresholds — with wrong answers that are plausible neighboring values, not nonsense.
- The scenario trap. "What should you do?" questions where one option is simply the most responsible are filler. A real scenario question makes every option sound responsible and only the code can break the tie.
What we changed
We found this problem in our own question bank — measuring it is how we learned everything above. So we rebuilt the pipeline around beating it: every question we generate now has to survive that blind evaluator before it reaches a student. If the question can be answered without the source material, it gets rejected and rewritten until only book knowledge cracks it. Our calculation questions — electrical load calcs, business math — sit at roughly 0% guessable, because you can't sweet-talk your way to a computed answer. We also keep a running ledger of what our own audits catch and throw out — it's published here.
The goal isn't to make practice feel hard. It's to make your practice score mean something, so the number you see the night before your exam is the number you'll see at the testing center.
If you're studying for a contractor, trade, or real estate exam, try a free practice preview on the exam page for your state — no signup required. And whatever you study with, run the five checks above on it first. Your license is too expensive to trust to questions a roofer can pass on vibes.