An AI coach should have to prove it's working
Cody BolithoJuly 20, 202612 min read

Ask your golf app a simple question: what did you tell me three weeks ago, and did it help?
Most apps can't answer the first half. None that I've used can answer the second. The advice gets generated, you read it, and it evaporates. Next session, fresh advice, no memory that the old advice ever existed. A human coach who worked that way would get fired by Christmas.
I've been building OpenCaddie around a rule that sounded obvious and turned out to be rare: the coach has to keep score on himself. But keeping score only means something if the advice was grounded in the first place. So there are really two obligations here, and most golf AI ducks both. Say where the advice came from. Then check whether it worked.
This piece is the long version of both. It's the most detail I've published about what sits under Mully, partly because people should be able to audit a coach, and partly because the name on the door is OpenCaddie and it would be a bit rich to be cagey about it.
Part one: what Mully is allowed to know
The first rule in the knowledge base is a restriction rather than a capability. Every benchmark Mully can reference traces to a named, checkable authority. If I can't attach a real source to a number, the number doesn't go in.
That sounds like table stakes. In golf content it very much isn't. The internet is saturated with confident numbers that came from nowhere: a swing tip that went viral, a forum post quoting a figure someone half-remembered from a fitting, a YouTube thumbnail promising the "correct" attack angle. A model trained on the open internet will happily reproduce all of it in a confident voice, and you'd have no way to tell which parts were real. So the knowledge base explicitly excludes social-media swing tips, influencer claims, and unverified forum numbers. A few figures I wanted got left out entirely because the only versions I could find lived inside an image or behind a paywalled PDF and I couldn't verify them properly. Leaving a gap is better than filling it with something that sounds authoritative.
Three tiers, because not all numbers are the same kind of claim
The sources are grouped into tiers, and the tier decides what a source is allowed to back.
Primary covers instrument makers, governing bodies, and peer-style research. TrackMan for ball-flight physics and launch data, Mark Broadie's strokes-gained work for how scoring actually decomposes, the USGA and R&A for rules and equipment standards. Anything that's a physical or numeric fact about a golf ball cites from here.
Dataset covers the large shot-tracking companies, principally Arccos and Shot Scope, who have millions of real amateur rounds. This tier backs the population questions: what a 15-handicap's greens-in-regulation rate actually looks like, how far people in your bracket really drive it, how often a 20-handicap three-putts. These are the numbers that let a diagnosis be about you rather than about a tour player.
Instruction covers credentialed teaching, and it's the tier I added last and thought hardest about. Drill methodology and feel cues have to come from somewhere, and the honest options are either "I made this up" or "here are the PGA-credentialed coaches whose method this follows." I went with the second. The instruction sources are named in the product, and they're teachers with real credentials and launch-monitor or 3D-capture-validated methods rather than viral-clip merchants.
Separating them matters because these are genuinely different types of claim. "Face angle largely determines start direction" is physics. "Here's what share of greens a scratch golfer actually hits" is a population statistic from a dataset. "Here's a towel drill for a low point that's behind the ball" is pedagogy. Blending all three into one undifferentiated voice is how coaching products end up sounding authoritative about everything and accountable for nothing.
What a benchmark looks like when it's actually doing work
Here's a real one. TrackMan publishes both tour averages and their Combine averages for amateurs, so driver spin has a proper ladder: PGA Tour driver spin averages about 2,545 rpm, while the average male amateur around a 14.5 handicap sits near 3,275 rpm, with a 10-handicap around 3,192 (TrackMan).
That ladder is the whole point. If your driver spins 3,400, an unsourced AI might tell you that's "high" and prescribe something. Against the right rung of the ladder, 3,400 is a completely normal number for a mid-handicap, mildly costly in carry, and probably not the biggest leak in your bag. Comparing you to the right cohort is what stops a coach from generating urgent-sounding advice about a non-problem. It's also why the benchmarks are segmented by handicap rather than served as one universal "good" number, which is the single most common way golf apps mislead people.
The diagnostic spine is ball flight
Mully diagnoses from what the ball did, then works backward to the cause. That ordering follows how ball flight actually works: start direction is governed primarily by face angle at impact, and curvature comes from the face's relationship to the swing path rather than from path alone (TrackMan).
Encoding those laws means the coach can name a cause instead of describing a symptom. A slice is an effect. The cause lives in face-to-path, club path, and where you're striking the face, and those three have to be reconciled against each other before anything gets prescribed. Strike location is part of that read, because a consistent toe or heel pattern produces curvature through gear effect that the path alone would never explain. Prescribe a path fix for what is actually a strike problem and you'll spend six weeks getting worse.
This is also why I'm unbothered by the swing-camera products. Grading a swing in isolation and then guessing at outcomes runs the causality backwards from how coaches actually work. Knowing your shoulder tilt was 38 degrees doesn't tell you your 7-iron leaks right when the pin is tucked.
The boring parts nobody markets
A few of the least glamorous pieces are the ones that keep the whole thing honest.
There's exactly one handedness rule, written once. Every directional sign, spin axis, club path, face-to-path, miss pattern, and drill cue mirrors through it for a left-handed player. That sounds trivial until you learn it wasn't always the case here: the rule used to be duplicated across separate parts of the system, the copies had quietly drifted, and one of them was missing entirely. A left-handed golfer was getting a subtly wrong read. Consolidating it was unglamorous work that fixed a real bug for real people.
Units get handled the same way. If you work in meters, every benchmark converts on the way out, so you're never mentally translating a yardage while reading your own diagnosis.
And the knowledge base carries a version stamp. When I revise the benchmarks or the coach's instructions, the version changes, so I can compare the quality of what came out before and after rather than going on vibes about whether it "feels smarter." That matters more than it sounds. Without a version, every change to an AI system is an untested opinion.
Deterministic analytics, AI for explanation
This is the architectural line that does the most work, and it's the answer to the obvious question of why an AI coach doesn't just hallucinate numbers at you.
The math is not the model's job. Your averages, dispersions, gaps, trends, and strokes-gained style reads are computed in code, deterministically, the same way every time. The model's job is to interpret that output, compare it against the sourced benchmark, explain it in English, and prescribe from a fixed catalogue of drills. It is instructed, in the strongest terms the format allows, never to invent data, numbers, or sources that weren't given to it. If a metric is missing, it works with what's there and says so.
So when Mully tells you your carry spread is 11 yards, no language model estimated that. Code measured it. The model is explaining a number it was handed, against a benchmark someone published, and prescribing from a catalogue a human wrote. That's a much narrower job than "AI golf coach" implies, and the narrowness is the feature.
Part two: checking whether any of it worked
Grounding gets you advice worth reading. It says nothing about whether the advice helped. That's the second obligation, and it's the newer half of the system.
Every drill Mully prescribes now becomes a tracked record. The record knows which metric the drill was aimed at, say 7-iron carry consistency or driver spin, when it was prescribed, and why. A nightly pass compares each open prescription against a deterministic read of your data and settles it with one of four outcomes: cleared, improved, no effect, or regressed.
Your next review opens by settling the last one. If the previous review told you to watch your low point and run the towel drill, the new review starts with what your low point actually did, and the drill gets marked accordingly. Cleared drills get retired. Drills that did nothing get replaced, and Mully says so in plain terms.
Why the measuring half is the hard half
Judging whether a golf number improved is easy to fake and surprisingly easy to get wrong even in good faith. Golf data wobbles. A warm afternoon adds carry. A range session where you only hit your favourite club flatters everything.
So the rules underneath are deliberately boring math. Each tracked metric has a noise floor, a minimum move that counts as movement at all. For driver spin that's 150 rpm. For greens in regulation it's five percentage points. If your own baseline is scattered, the floor rises to half your spread, so a noisy player needs a bigger move than a consistent one before anything registers.
Then there's the gate I care most about: nothing gets marked cleared off a single good session. Cleared means you were behind the benchmark for your handicap, you're now at or past it, the move beat the noise floor, and your last two data points both held there. A single flushed session proves you had a good session. Faults only count as fixed once the improvement sticks.
The full rule set, including the actual thresholds and the benchmark tables by handicap, is public at opencaddie.ai/learn/how-mully-measures-progress. That page renders from the same constants the product runs, so if I tune a threshold the page changes with it. I'd rather show the mechanism than ask you to trust the adjective.
What a drill can never be credited for
Here's the part I want to be strict about, because this is where coaching products usually start lying by implication.
When a drill gets marked cleared, the honest reading is: the number crossed your benchmark and held while you were running that drill. Maybe the drill did it. Maybe you also played three extra rounds that fortnight, or the weather turned, or you finally stopped using a worn glove. The evidence stored with each outcome records what the number did and when. The language stays at "cleared while you ran it," and Mully is explicitly forbidden from claiming the drill caused the change.
Same discipline on the aggregate side. I don't have the volume yet to say "this drill clears low-point faults for most players like you," so Mully doesn't say anything of the sort. If that kind of claim ever shows up, it will be because the data earned it.
The gaps, stated plainly
Since the point of this piece is auditability, here's what the system can't do yet.
It can't establish causation for an individual, and by design it won't pretend to. It can't make population claims about which drills work best, because that needs far more players and far more settled outcomes than I have. Some per-club benchmarks are interpolated along a known gradient rather than measured directly, because the published tables don't cover every club, and those are marked as estimates rather than dressed up as measurements. And a metric that your equipment doesn't capture simply isn't part of your read, which is why a diagnosis from a basic launch monitor is thinner than one from a full unit.
None of that is comfortable to write on a marketing site. All of it is true, and a coach that hides its limits is a coach you can't calibrate.
What I publish and what I don't
Being called OpenCaddie sets an expectation, so let me draw the line where it actually sits.
Published: the source list and what each source backs, the benchmark tables by handicap, the ball-flight laws, the progress thresholds and noise floors, the four outcome states, and the sourcing policy itself. The pages that show this render from the same constants the product runs on, so they can't quietly drift from reality. You can read the sourcing tree at opencaddie.ai/learn/sources and the measurement rules at the link above.
Not published: how the context handed to the model gets assembled and prioritised, how drills get selected and ranked against a given fault, and the prompt engineering itself. That's the part I've spent a year on, and publishing it would be handing a competitor the build rather than giving a golfer transparency. The distinction I'm drawing is between the evidence and the machinery. You're entitled to audit what grounds your advice and how progress is judged. You don't need my prompt files to do that.
If that sounds like a convenient place to draw a line, fair. Judge it by whether the published half is enough to check my work. I think it is, and it's a great deal more than "our AI analyses your swing."
Why bother with all this
Because a prescription that gets checked changes the coaching that follows it. Once outcomes exist, the next review can't re-prescribe a drill for something that already cleared, and it has to own the prescriptions that went nowhere. The work you get handed in August is built on what measurably happened in July, and the whole thing compounds instead of resetting every session.
It also changes my incentives, which is maybe the realer point. A coach that publishes its own sources and its own hit rate has to care about being right more than sounding right. Every one of those published constants is a hostage. If a threshold is badly chosen, it's sitting on a public page with my name on it.
Whatever you use for your golf, an app, a teaching pro, a buddy with opinions, it's worth asking the two questions I started with. Where did that come from, and did it work? The answers should be specific, and somebody should have written them down.
Sources
Every benchmark above traces to a named source. Here they are.
OpenCaddie
Coaching, not another dashboard
Drop in your TrackMan, SkyTrak, Foresight, or GSPro export and OpenCaddie reads it, names the fix, and prescribes the drill.
Try OpenCaddie free