On September 18 someone bought our Texas Residential Appliance Installer practice exam. At 3:30 the next morning, before anyone on our side was awake, an automated audit went through all 319 questions in that pool with no book, tried to answer each one three separate times, and had to explain how it got each answer. It pulled 73 questions — 23% of the pool — that could be answered without knowing the code. By that evening, 33 replacements were being written from the sections of the source book that no question tested yet.
The customer never saw a ticket, a notice, or a "pardon our dust" banner. That is what "we audit our own exams" means in practice, and this post is the ledger: what the audits found, what we threw out, what we put back when we were wrong, and where we still lose.
Why we audit ourselves at all
In August we ran our blind-solve test on our whole catalog — 3,054 questions sampled across every live exam. The evaluator gets no study material and has to show its work: did it know the fact, or did it crack the question through wording, common sense, or elimination?
The number was 27% guessable. Our electrical calculation exams sat at 0% — you can't sweet-talk a load calc. Our real estate exams sat at 52–65%, level with the softest commercial prep book we've measured, and a long way from the 10% a major test provider's own prep book scores on the same rubric.
Two other things fell out of that audit:
- Our own difficulty labels were noise. Questions we had tagged "easy" were 31% guessable, "medium" 26%, "hard" 25%. Six points of spread. The label was steering which questions went into a test while telling us almost nothing about whether the question tested knowledge. We stopped trusting it.
- The problem was our distractors, not our facts. On substance our questions scored above both commercial books. The keyed answer was just the longest, most careful option, and the wrong answers disqualified themselves. Ten mechanical rules came out of that — equal option lengths, no keyword echo from the stem, no self-defeating distractors, absolutes in zero options or at least three — and they now sit in the generator's instructions and its reject gate, in front of every question it writes.
The audit, and the human check on the audit
Here is the loop that runs today. Every exam someone has paid for is re-audited whenever its question pool changes. The blind solver attacks every question in the pool. Anything it solves gets judged twice more, and only a 2-of-3 confirmation counts. Confirmed questions are benched with the reason attached — they are not deleted, and they are never quietly put back.
That sounds airtight. It wasn't, and the person who proved it was a roofer.
The solver reports how it solved each question: a wording tell, plain common sense, or a chain of "logic" from the stem. We had been benching all three. So we handed one of our founders — a roofing contractor with no electrical training — 60 electrical questions, half of them the solver had cracked by "logic" and half it could only answer by knowing the code. No book, no lookup.
He scored 11 of 30 on each half. Identical. What the machine called "logic" was electrical knowledge a layperson doesn't have. We had been throwing out questions that were actually fine.
So we changed the rule — only tells and common sense count as leaks now — and restored 1,364 questions across 34 exams that had been benched on that route alone. 1,277 stayed benched. Being wrong in the strict direction is the cheap kind of wrong, but it was still wrong, and the fix came from a human test, not from the model grading itself.
The leak the solver couldn't see
The same 60-question sheet exposed something no per-question audit could catch: the correct answer was B far too often. Across the fleet, 37% of keyed answers sat in position B against a 25% baseline; on our oldest questions it was over 50%. Someone who answered B on every question of that roofer's sheet would have scored 40% — better than the roofer, who read all sixty.
That is a free lift for a test-taker and a fake lift for a student, and it existed in every practice test we served. On September 4 we re-shuffled 5,536 questions so every letter now holds the answer about a quarter of the time, remapped the Spanish translations and 2,379 already-answered session records so old results still point at the right text, and kept an undo log. Every guessability number we'd quoted before that date was measured on the skewed corpus, which is why we're restating them here.
The ledger
Every one of these is a change to exams people had already paid for.
- Real estate, August 28: about 460 questions retired across six exams in one sweep, then regenerated under the new distractor rules.
- Four exams, September 1: the first nightly audit benched 416 questions in one night. The three electrical pools in that batch were flagging at 41–46%; after replacement they re-audited at 13–17%.
- Wrong code edition, August 29: Oklahoma's mechanical exams test the 2018 mechanical and fuel gas codes. Over 120 of our live questions cited the 2021 editions, and a 2021-keyed answer can be wrong under 2018. Benched, book re-ingested, regenerated. A fleet sweep afterward found zero.
- Duplicates, September 16: twenty exams checked for pairs of questions that test the same fact with different words. 2,563 questions rejected and 1,788 shared links cut. A practice test that asks you the same thing twice is a shorter test than it looks.
- Stale bulletins, September 17: twenty exams were built from a candidate bulletin the state had since reissued. One of our pages said the passing score was 75% when the state's own table says 53 correct of 74 — about 72%. Two states had raised their fees. All twenty re-pointed, copy corrected, and every exam now sits on a weekly watch that flags the next reissue.
- Home improvement, September 19: our own Maryland pool measured 67% guessable — the softest we have ever recorded. 55 questions benched that afternoon, replacements queued.
Where we win, where we don't
We also run the same solver on purchased prep material, because "we audit ourselves" means nothing without a yardstick. No names — the pattern is what matters.
- Illinois roofing: a purchased vendor set judged 48% guessable, 17% of it by pure wording tells. Ours: 30% and 4.7%. Their questions were mostly short fill-in-the-blank stems; ours are scenarios you have to read.
- Oklahoma business and law: the vendor set and ours came in dead even at 55%. Business-law questions are hard to write tightly, and we are not there yet on that exam. That number is in this post because leaving it out would be the kind of thing this post is against.
What "restructure" actually means
Benched questions go to a review pile with the reason attached. Replacements aren't "20 more questions on this topic" — they are written against the gap list, the concepts in the mapped source book that no surviving question tests, so the pool gets wider instead of just deeper. New questions face the same blind solver before they ship, then the pool is re-audited. We hold every topic at four to five times its weight on the real exam, so a full-length practice test rarely repeats a question. And the moment the pool changes, that exam goes back on the nightly list.
What we can't claim yet
One customer took 391 questions with us over six days, went from 60% to 88% on practice finals, and passed his state exam. We would love to say we did that. We don't know that we did — he may have studied with other things too — so every exit survey now asks what else you used before we count anything as ours.
What to do with this
Ask any practice source three questions: when did they last audit their own questions, what share can be answered without the book, and what did they throw out. Then run the five two-minute checks from our first post on a handful of its questions. If the answer to all three is silence, the practice score it gives you is a number, not a measurement.
If you want to see what a pool looks like after this treatment, every exam page on the site has a free preview — no signup. Take it cold, then ask yourself whether you knew the answers or just recognized the shape.