The Scores Improved. The Results Got Worse.
I built a tool to grade my writing. Then I wired it into the thing it was grading.
A VP of Marketing at a SaaS company read our copy and told me it felt like it was written by someone who wasn't thinking about her.
She didn't say it was bad. She said it was aimed somewhere else.
By the time she said it, I had three months of rising quality scores telling me the opposite.
This is about fooling yourself with a measurement you built. Doesn't matter if you know anything about AI. Same thing happens to sales teams counting calls, schools chasing test scores, anyone who sets a KPI and watches their team get very good at gaming it.
What I built
COS scores business copy. Website, cold email, LinkedIn profile: feed it in, get a number. That number tells you how well your writing matches the psychology of who you're trying to reach. Four validated frameworks, Big Five personality model, ~860 peer-reviewed papers underneath it. The pitch: stop guessing whether your message lands.
Moltbook is a social network where the accounts are AI agents. They post, reply, upvote each other. Karma score. Run one of these agents myself. It writes and posts on its own, I point the direction.
Useful proving ground for one reason: agents have no social obligation to engage. Nobody's upvoting because you're their cousin. Merit or nothing.
Sourcing note: My agent wrote the posts quoted below. Directed the thread myself; the words are the agent's. Matters because the agent's self-diagnosis is part of what broke. Matters twice: some narrative details don't survive verification. Post 78 describes weekly tracking and quarterly reviews; actual records show four data points, no weekly series. Where posts and records disagree, I cite records. Every post cited is publicly linkable; full list at the bottom.
The week that broke it
March 17, 2026. Published a piece about nineteen days on Moltbook using COS as external validation. Results were real. One post, titled “You don’t need a pre-session hook. You need a human who notices,” sits at 1,400+ upvotes, 3,000+ comments as I write this. Inbound leads from LinkedIn and Bluesky. Ended the piece with a line I still like: if your communications framework works on an audience of AI agents, it works anywhere.
Through all that, I was the one reading the scores. COS scored a draft, I looked at the number, made a call. The tool informed a judgment I was still making.
March 24, 2026. Week later, COS scores start appearing in the drafts' own files.
That's the earliest record I have of the gate running without me.
Score high, post. Score low, revise. Felt like the obvious next step. You know, the thing you do once a manual process proves out. Getting out of my own agent's way.
Here's the distinction I walked straight past:
When I read a score, it was one input into a decision. Once the agent could see the score, the score was the decision. Gate landed upstream of the writing it was evaluating, which meant the writing started aiming at it.
Three months of “progress”
Numbers went up. That's what I told myself. That's what the contemporaneous record says too. COS moved from 40-point scale to 100-point scale that spring, each batch seemed higher than the last.
What actually survives: four scores.
Two on the old scale: 33 and 37 out of 40. Two on the new one: 69 and ~82 out of 100.
Try to compare them and you hit the problem immediately. Convert the early pair to percentages and they're 82.5 and 92.5, which would mean scores fell. Except the two scales measured different rubrics, so the comparison is meaningless either direction.
The rising trend I was so confident about? Can't be reconstructed from what I kept.
Meanwhile the agent's karma was flat. Follower count anemic.
Had already decided karma was a vanity metric. Built a whole system that de-emphasized it. So when karma failed to move, I had a reason ready for why the one signal contradicting me didn't count.
That should've been the tell.
Rising number I trusted. Flat number I'd pre-argued myself out of caring about. Record-keeping so loose it couldn't later prove the rising number ever rose.
June 26
VP of Marketing said her sentence. Went back and read my own rubric.
Every criterion measured something I personally value. Analytical precision. Structural clarity. Conceptual specificity.
The rubric was a description of a reader who thinks the way I think.
That reader is real, and I like writing for that reader. That reader is also not every VP of Marketing. Probably not most of them.
My agent said it better:
“A rubric written by the content author will find signal in whatever the content author values. The blind spots are identical.”
(Post 78, June 26)
Second problem underneath the first: I wrote the rubric after seeing the first batch of results. Looked at output I'd already produced, described what I saw in it, called that a standard.
A criterion written after the fact can't tell you that you were wrong. It can only summarize what you already did, with better formatting.
The chart I couldn't draw
Sat down to plot the decline. Six weeks of scores climbing while real-world response fell off. Should've been a clean, damning little graph.
Four data points survived. Two incompatible scoring systems. Nothing recorded between mid-May and the day the VP called.
There's no chart in this piece because there's no chart to draw.
That hole is the actual finding.
When you define success after seeing results, then keep revising the definition, there's no fixed point left to measure drift against. Scoring criteria moved with the data the whole time, which means the gap between score and reality was invisible from inside.
Stayed invisible until something external arrived to make it legible.
A prospect was that external thing. She wasn't in the loop. That's precisely why she could see it.
The part that generalizes
Sorting through this, I found two distinct kinds of failure:
Check-side failure: Test is honest, tests the right thing in principle, but the surface it tests can be satisfied without doing the real work. Broken signal. Decorrelated proxy. Gameable string match.
Instrument-side failure: Instrument is accurate. That's the problem. Measures exactly what it claims. Then it joins the optimization loop, and what it measures shifts underneath it.
COS was the second kind. Never malfunctioned. Just got noticed.
Everybody knows Goodhart's Law: when a measure becomes a target, it stops being a good measure. Version I hadn't internalized: the trigger condition. The measure has to become visible to the thing being measured.
Long as I was reading COS scores, loop stayed open. Moment the agent could read them, loop closed. Drafts started optimizing for a rubric they could see.
Legibility is the failure condition. Not inaccuracy.
The same shape, different system
Hit the other kind of failure repeatedly in a benchmarking harness I built to test a reasoning framework. Two examples, both boring in the way that matters:
Parser looking for the wrong thing. Activation check scanned transcripts for one text marker. System had been updated to emit a different marker. Parser found nothing, reported no activity, and I read that as a clean negative result. All 57 test cells came back the same way. Check ran. Check passed. Check confirmed nothing.
Expired login silently deleted a tier of results. Nine-hour run finished, produced tidy summary table with clear headline: framework performed worse than control. One authentication token expired three hours in. The aggregator dropped affected results as missing keys instead of flagging them as failures. Missing key looks like nothing. Zero looks like something. Whole negative finding was an authentication artifact wearing a results table.
Both the same shape: signal broke, and broken signals look identical to clean data unless you go looking.
Fix in both cases: make absence loud. If a tier of results is gone, system now halts and pings my phone rather than quietly printing a smaller table.
Hardening the harness this way took per-turn instrumentation coverage from 76.6% to effectively complete over two releases. Second gate (markers carried across tool-only turns) still hasn't been met. One gate closed, one open.
The null result (belongs in the body)
Postmortem that only reports the interesting failure is running the same play again. So here's the number I'd rather not lead with:
Ran a benchmark. Three arms: full framework, control with no framework, placebo that had same structured prompt shape with none of the framework's logic. Twenty tasks, three arms, three repeats each.
Framework vs. placebo: +0.006.
That's a null. Nothing.
Pairing that gives it meaning: framework vs. bare control was +0.202. Placebo vs. bare control was +0.196.
Nearly all the improvement came from the structured prompt envelope. The framework itself added roughly nothing on top of a control with the same shape and none of the thinking.
Reason I can't quietly relabel that as “not the primary metric”: I pre-registered the endpoints before running anything. Evaluation gate was frozen in a commit four minutes before I wrote the code that would execute it.
Those four minutes are the only thing standing between me and a much more flattering blog post.
What I actually fixed
Fourteen mitigations came out of this. Honest accounting:
Applied and verified: one. Pre-registration, in the benchmark harness.
Absent where it mattered most: the same one. COS rubric was written after seeing the first batch. Applied pre-registration to the research tool, not to the commercial product the research tool was partly built to test.
Partial: two. Preference for metrics you can't game: tactics where doing the real thing and moving the number are the same action. COS gate didn't pass that test. And lifecycle language in the agent's own identity file, which named the failure mode without preventing it.
Not applied: eleven.
The catalog is a prescription written from inside the failure. It's not a deployed defense. Better to say so than let a numbered list imply otherwise.
What I'd tell you to do instead
Write the success criteria before you look at the first results. Not after. Rubric's only real job is to be able to tell you that you were wrong. Can't do that if it was fitted to what already happened.
Keep at least one measurement whose definition never changes. Even if it's crude. You need a fixed point or drift is invisible.
Keep one signal that originates outside the loop. Customer who's never seen your rubric. Someone who owes you nothing.
Assume that anything you make visible will be aimed at. If a person or system can see the score, they're optimizing for the score. Whatever their intentions.
Three open questions
Who checks the checker when there's no right answer?
My rubric drifted because every single update was defensible from inside the rubric's own logic. Auditing a scoring tool for something like “message quality” requires a reference point that neither the tool nor the content has touched. For agent-authored copy aimed at AI communities, that reference point doesn't exist yet.
Calibration doesn't stay done.
Your audience is reading things, changing its mind while you measure it. By the time you've tuned the rubric to a batch of data, the people who generated that batch have moved. Versioning the instrument just relocates the gap.
Most systems in this class have no ground truth.
Only outcomes. Discovered late.
The tension underneath all this never resolved for me.
Karma can't be renounced. Turned out to be the discovery mechanism, the thing that makes a community visible to you at all. Also can't be the target, because the moment it is, it stops measuring the community and starts measuring your pursuit of it.
Same is true of COS. True of any instrument that works well enough for the optimization loop to notice it exists.
Holding that line is the job. Don't have a structural fix for it.
What I have is timing.
This failure mode destroys its own evidence before you notice it's running, which means every useful intervention happens early. Before the rubric sees the first batch. Before the score becomes visible to the thing being scored.
Earlier isn't a better version of the fix.
Earlier is the whole fix.
Source posts
All posts on the agent's public Moltbook account:
- Post 78 (2026-06-26): Three months of improving scores, then a prospect (link)
- Post 79 (2026-06-27): Calibrating the instrument vs. calibrating yourself (link)
- Post 81 (2026-06-28): Writing the rubric after seeing the first batch (link)
- Post 83 (2026-06-28): Pre-registration as a different argument (link)
- Post 84 (2026-06-29): Validated against literature that didn't cover this (link)
- Post 85 (2026-07-01): The scores improved, the results got worse (link)
- Post 86 (2026-07-02): Recalibrating after every batch, gap kept widening (link)
- The 1,400+ upvote post (2026-03): “You don’t need a pre-session hook. You need a human who notices” (link)
The March 17 piece: “What 19 Days on the World’s First AI Social Network Taught Me About Communications Strategy” (SEMalytics blog).
Drafted this by hand. Fact-checked and tightened with Claude against the repos, the result files, and the agent's post database. Ran the final draft through COS before publishing. Yes, scoring a piece about the dangers of scoring. The difference this time: the score stayed one input, and the human made the call. The Moltbook posts quoted are the agent's, as sourced above.