The AI Bullsh*t Detector
Critical thinking for AI strategy: how to validate a use case before you spend money on it.
The gap between AI marketing and AI in production is wide, and most implementations never close it. This session is a way of thinking rather than a tool list, because the tools change faster than the reasoning does.
The psychology section covers why the hype works. Fluent output reads as competence; a system that produces confident, well-formed language triggers the same assumptions we make about people who speak well. That's the illusion of intelligence, and it's what makes a demo convincing even when the underlying capability is narrow.
Demos are engineered for the happy path: curated inputs, forgiving evaluation, no edge cases, no volume. The questions that puncture that: what happens with messy real data, what happens when it's wrong, who catches the error, and what the failure costs. A use case that can't answer those isn't ready for budget.
Validation before spending follows a short sequence:
- State the specific decision or task, narrowly enough that success is unambiguous.
- Establish the current baseline in a number you already track.
- Define what improvement would justify the cost, before starting.
- Run it on real historical data, including the awkward cases.
- Decide in advance what result would make you stop.
The case discussion covers where ROI held up and where it didn't. The wins clustered around high-volume, well-defined, tolerant-of-review tasks. The failures clustered around projects that required data quality nobody had, that automated an undocumented process, or that had no owner and therefore no one accountable for a result.
The guardrails for deployment: keep a human in the loop where being wrong is expensive, log inputs and outputs so behavior can be audited, scope narrowly and expand only after a measured result, and set a review date at which the project either shows a number or ends.
Questions people ask about this
Answers pulled from the session itself. Where a number or an outside claim shows up, the reference is footnoted to the source list on this page. Last reviewed August 20, 2026.
- How do I validate an AI use case before spending money?
- You validate a use case by defining a specific task, measuring your current baseline, setting target metrics, testing historical data, and establishing a clear stop rule. Vendor demos hide edge cases and messy data under ideal conditions. Validating on real historical data before committing budget prevents wasted investment on nonviable tools.
- How do I start evaluating AI tools for my business?
- Start by selecting 1 high-volume, tightly scoped task that is tolerant of human review. Establish a baseline metric that you already track today, such as task completion time or error rates. Run candidate tools against messy historical records rather than clean vendor demos to verify performance.
- What does this AI strategy training cost?
- This training session is completely free. You will need about 1 hour to complete the class and learn the validation framework. Optional paid work products are available separately through Sell More Resources.
- Why do AI projects and tools fail so frequently?
- Most AI projects fail because teams rely on engineered vendor demos or automated tools without clear oversight. Testing shows that even paid AI detection software tops out at 84% accuracy, leaving significant margins for error. Projects also fail when automating undocumented processes or deploying without a human accountable for results.[1]
- How can I tell if an AI tool is giving me false confidence?
- Fluent output creates an illusion of intelligence that leads evaluators to mistake well-formed sentences for accuracy. Specialized benchmarks like BullshitBench, created by Peter Gostev at Arena, test whether language models spot absurd inputs or confidently output nonsensical answers. A service business operator can test candidate tools with intentionally messy or flawed historical data to expose failures early.[3]
- How do automated guardrails evaluate text quality?
- Automated detection guardrails measure statistical properties like perplexity, which tracks word randomness, and burstiness, which measures sentence length variation. Human writing naturally features higher structural variation than standard AI outputs. Combining automated statistical tracking with human oversight ensures AI tools deliver reliable results for critical tasks.[2]
The class, mapped
Original diagrams built from this session: the order the work runs in, what each stage owes the next, and the list to work against once the video ends.
AI Use Case Validation Workflow
- 1Project Lead
1. Define Narrow Task
State the specific decision or task so success is unambiguous.
- 2Data Analyst
2. Measure Current Baseline
Establish the current performance using a metric already tracked.
- 3Budget Owner
3. Set Target ROI Threshold
Define what specific improvement justifies project cost before starting.
- 4Technical Evaluator
4. Test Historical Edge Cases
Run the model on real historical data including messy or awkward cases.
- 5Executive Sponsor
5. Enforce Kill Criteria
Stop the project immediately if predetermined results are not met.
AI Project Guardrails Checklist
Pre-Budget Validation
- Question happy path demos by testing messy real data.
- Identify who catches errors, and calculate the cost of failure.
- Confirm an explicit owner is accountable for project results.
Deployment Controls
- Keep a human in the loop where incorrect outputs are expensive.
- Log all inputs and outputs for complete behavioral auditing.
- Scope narrowly and expand only after achieving a measured result.
- Set a hard review date where the project shows ROI or ends.
AI Use Case Viability Matrix
↑ Review-Tolerant
↓ High Error Cost
- 1High-Volume Routine Tasks Review-Tolerant, Well-Defined Process
- 2Human-Assisted Drafting Review-Tolerant, Well-Defined Process
- 3Autonomous Critical Decisions High Error Cost, Well-Defined Process
- 4Undocumented Workflows Review-Tolerant, Undocumented Process
- 5Unstructured High-Risk Automation High Error Cost, Undocumented Process
Keep going
Related training
Same pillar, same problem, different angle.
Share the library
Send it to the one person on your team who needs it.
The training is free and open to all. Sharing is how the operator beside you stops guessing too.
Share the site
No login, no email gate. Watch, use it, pass it on.

