AI in practice
We Audited 40 AI Finance Skills. 4 Were Theater. Here's What Actually Works.
The audit in one glance: 36 working engines, 4 retired as theater.
A few months ago I did something that felt a little like auditing my own books: I ran every AI skill we've built for construction and SMB finance through a formal audit, and I killed a tenth of them.
Not because they were broken. Because they were theater — they looked like capability, and they did nothing a blank chat window couldn't do.
This is the part of the AI conversation nobody wants to have in public, because it's awkward. The skills were well-written. They had nice names and clear instructions. And when I asked the hard question — "does this beat just asking the model directly?" — four of them had no answer.
If you sell AI tools, or you're thinking about buying them, you should know how common this is. And you should know the three questions that separate the tools that work from the ones that just look like they work.
The number that should worry you
This isn't just me being hard on myself. There's now actual research on this, and it's worse than my 4-out-of-40.
A benchmark called SkillsBench (published earlier this year) was the first serious attempt to measure whether AI "skills" — pre-packaged instructions that teach a model to do a specific task — actually improve results, versus whether the model could've done it alone. The finding that stuck with me: practitioners have had no reliable way to know whether a given skill helps or just adds noise. Most of what's out there was never measured at all.
Then there's the study with the blunt title "Skill Use or Skill Theater?" The researchers found that the obvious ways to tell if a skill is working — did the model mention it? does the output look like the skill? does an AI judge think it helped? — don't actually tell you. A skill can be "used" and have zero effect on the answer.
And my favorite, because it's so counterintuitive: a study on skill libraries found that as you grow from 5 skills to 100, the system's ability to even pick the right skill collapses from 30% accuracy to 3%. A bigger library is a worse library. The companies bragging about their "1,000+ AI skills" are, measurably, making their product less reliable.
So when I tell you I retired 4 of our 40, the honest framing is: I got off easy. Most people selling this stuff haven't looked.
The 4 that failed — and why
I'll name the pattern, because it's the whole lesson.
The four I killed were things like a "guest recovery playbook" for restaurants and a "founder time reclaim" framework. They were advice. Good advice, even. But when I ran them against a bare model with no skill loaded, the model produced essentially the same output. The skill was just a prompt-shaped essay. It encoded nothing the model didn't already know.
Everything that survived had one thing in common: a deterministic engine. Real code that computes something — a WIP schedule, a retainage ledger, a three-way invoice match — with rules a blank model would get wrong or invent. The skill wasn't there to make the model sound smarter. It was there to make the model compute correctly where correctness is the whole job.
That's the line. On one side: prompts that flatter the model. On the other: engines that do math the model can't be trusted to do from memory.
The three questions (steal these)
This is the exact protocol I used. If you're evaluating any AI tool — including mine — run it through these:
1. Does it encode something the model couldn't know?
Private, domain-specific knowledge: the actual formula for retainage, the AIA G702/G703 structure, your state's statutory lien-waiver deadlines, prime-cost thresholds. If the "expertise" is generic — "communicate clearly," "follow up promptly" — the model already has it. You're paying for a wrapper.
2. Does it demonstrably beat the bare model?
Not "does the output look good." Does it measurably beat asking the model directly, on a real task? If the vendor can't show you that comparison, they haven't run it. We now keep golden fixtures — known-right answers — and test that our tools reproduce them exactly. If a tool can't show you its known-right answer, be suspicious.
3. Is it permanent logic, or a patch for a model quirk?
Some skills exist only to work around a limitation today's model has. Those have a shelf life — the next model release absorbs them and they become maintenance drag. The ones worth keeping encode rules that don't change when the model does: tax thresholds, bond math, statutory forms.
If a tool fails question 1, it's theater. If it fails question 2, it's unproven. If it fails question 3, it's temporary.
Why I care about this more than most
In funds control, a wrong number isn't embarrassing — it's a breach. A retainage figure that's off by a dollar, a WIP schedule that overstates margin, a lien waiver dated wrong: these are the things that get a contractor's bonding pulled. I've sat across from the owner whose "profitable year" didn't survive a real review.
So when AI tools hand a contractor a confident, well-formatted, wrong number, it bothers me in a specific way. The formatting makes it worse, not better. It looks bondable. It isn't.
That's why our rule is that the free skill on GitHub and the paid product compute the same number to the penny, tested against golden fixtures. Not "about the same." The same. In a market full of AI vapor, verifiable is the only thing that should earn trust — especially from an industry that runs on surety bonds and audited statements.
What actually will save you time
If you bought ChatGPT Team or Claude for your business and it's mostly generating emails you're not sending, here's the honest path:
- Stop collecting prompts. A hundred generic prompts are worse than five good ones — measurably. Pick the few workflows where a wrong answer costs you money.
- Automate the math, not the writing. The ROI is in the deterministic stuff: job costing, reconciliation, retainage tracking, cash-flow forecasting. Things with a right answer. Let the AI handle the judgment around the edges, not the arithmetic at the center.
- Keep a human on the money. Every number that flows into a bond capacity decision or a pay application should pass through a person. The tools that promise "no humans needed" are the ones that get your bonding pulled.
The goal isn't an AI that does your job. It's an AI that does the 85% of your job that's arithmetic and chasing paper, so you can do the 15% that's judgment. That 15% is the only part that was ever really yours.
Want an honest answer about your workflows?
Book a 20-minute operations audit. We'll look at how your week actually runs and tell you whether AI fits, or whether you'd be wasting your money.
Book the 20-minute audit