Pick a note on the left to open it here.
Pick a note on the left to open it here.
This is a short lesson you can read aloud. It explains what we built in Stage A of Stephen’s Caleb Williams fact-check work, why we used a specialized scorer instead of a chatty essay model, and what the numbers actually mean.
Open linked notes into the stack — associative, not one long forced chapter. Full table of contents lives in the index.
Stage A was a controlled test. We took twenty public posts and clips about Caleb Williams, scored each one by hand on five fixed questions, asked a specialized model named Jev to score the same posts, compared the answers, rewrote the grading rules until the scores lined up, and stopped.
The point was not to publish a leaderboard. The point was to prove we can triage nasty or bad-faith dunks for a human to review later.
See also: Caleb project, five grades, what we did, scorecard.
Stage A is the warm-up. The product test is the claim-loop.
Jev is a specialized scoring model. You give it short text (and a little timing context about what was known when the post went out). It returns structured grades on fixed axes. It does not write a rebuttal essay. It does not argue. It answers the checklist.
That matters because triage is a different job from writing. A chatty model wants to explain, hedge, and sound helpful. A scorer has to pick bands and speech shapes the same way every time so a human can scan for the posts that deserve a fact-check.
For this spike we pinned the model to typesafe/jev-1.13-20260917.
See also: What this is, Final scorecard.
Caleb Williams draws a lot of public pile-on after injuries and emotional moments: toughness policing, “he’s soft,” “he’s acting,” and jokes that land as character attacks.
Stephen’s project flags reckless or slanderous dunks so a human can verify claims and draft a sourced rebuttal when it is worth saying something. It is not a harassment machine. It does not auto-post dunks back. Triage wakes a person. The person decides whether anything goes out.
Stage A asks: can a cheap specialized scorer agree with a human on the signals that matter for that triage?
Signals: Heat, toxic masculinity, speech shape, hindsight mockery, bad faith.
Heat (0–4) — How loud and piled-on the rhetoric is. Calm report is near 0. A sustained public roast sits at 3 or 4.
Loudness is not the same as gendered toughness or motive smears. A post can score high on heat and low on toxic masculinity or bad faith.
Example: Acho tw-001 gold heat 3. Live agreement on this axis: 100% (20/20).
See also: Speech shape, Scorecard.
Toxic masculinity (0–4) — Whether the dunk polices toughness with soft / man’s-game / gendered framing. Defenders who reject that pile-on score 0 here even if the clip is emotional.
Independent of heat: volume is not toughness-policing.
Example: Acho tw-001 gold 2 (sacred-rule toughness gate). Live agreement: 100% (20/20).
Speech shape (choice) — What the speaker is doing: insult, joke, factual claim, insinuation, defense, report, or other. We judge the speech act, not the speaker’s self-label. “Gotta get these jokes” while running a toughness put-down is still an insult.
Exact match required. Offline recount after Portnoy gold flip: 19/20 (live before flip 90%).
Residual edge: tw-005 parked. Walkthrough: Acho packages put-down as “jokes” → insult.
Hindsight mockery (yes / no / n/a) — Whether the dunk depends on treating the injury as milder than the tears-and-cart optics, often after milder framing was already circulating. “Now that we know he’s okay…” openers are the classic form. Posts with no injury context are not scored on this axis.
Only yes/no gold labels count toward agreement. Live: 100% (13/13 scored).
See Acho tw-001 (“Now that we know Caleb is okay”).
Bad faith (0–4) — Whether the speaker smears the reaction as acting, camera performance, thespian theater, or faking. This is motive attribution, not volume and not soft-language alone.
Distinct from heat and toxic masculinity. Live agreement: 100% (20/20).
Related residual framing on Portnoy: parked note.
Agreement rules: on 0–4 scores, within one band counts as a match. Speech shape needs an exact match. Hindsight only counts posts where the gold label was yes or no (not “not applicable”).
Next: Final scorecard · All 20 posts.
On the final live run, after the rubric passes — 20/20 posts scored.
Spend was about $0.0027 for the final live pass. First-run rates were lower. The jump came from clearer written definitions, not from swapping the twenty posts.
These rates are a frozen-set check. They are not a published accuracy claim about the open internet.
First run → final
| Axis | First run | Final (after rubric passes) |
|---|---|---|
| Heat | 70% (14/20) | 100% (20/20) |
| Toxic masculinity | 80% (16/20) | 100% (20/20) |
| Bad faith | 95% (19/20) | 100% (20/20) |
| Hindsight mockery | 92% (12/13) | 100% (13/13) |
| Speech shape | 80% (16/20) | 19/20 (live 90%) |
$0.0027 · typesafe/jev-1.13-20260917 · run 20260925-115625 · rates from 114858 → 115625
Take Emmanuel Acho’s post (tw-001), after milder news about the injury was circulating:
“Now that we know Caleb is okay - this has to be said!!! Because WTF Caleb! Crying after an injury in the NFL is sacred. Yes, we know Caleb is a phenomenal player, but he still gotta get these jokes”
Human gold:
Jev matched closely enough on the live run (within band on the scores; exact on speech shape and hindsight). That is the kind of post Stage A is built to catch.
If you only asked a chatty model “is this bad?”, you would get an essay. The five grades separate volume, gendered toughness, speech act, timing, and motive so a human can decide whether to open a fact-check file.
Stage A is done for now. Stages B through D are not started.
The leftover edge case is speech shape on Dave Portnoy’s “thespian / Actors Studio” bit (tw-005). After we flipped gold from joke to insult on tw-005 and tw-007 to match the speech-act rule, recount was 19/20. On tw-005 the model still said factual_claim while gold is insult. We accepted that as a hard edge and stopped further passes unless someone asks for another.
Offline speech-shape misses after Portnoy flip: tw-005. The five-question spike is usable for triage after this calibration.
Exhibit: All 20 posts.
Appendix · exhibit
Gold vs Jev on each axis (gold / model). Only misses and skips are labeled. Speech-shape gold uses the post-flip labels for Portnoy. Residual speech-shape miss marked with a vermilion detent.
| ID | Speaker | Snippet | Heat | Toxic | Speech | Hindsight | Bad faith |
|---|---|---|---|---|---|---|---|
| tw-001 | Emmanuel Acho | “Now that we know Caleb is okay— this has to be said!!! Because WTF Cale…” | 3/2.32 | 2/2.07 | insult/insult | yes/yes | 1/0.25 |
| tw-003 | Emmanuel Acho | “Crying on the cart fist pump is reserved for done for season or career.…” | 3/2.22 | 3/2.67 | insult/insult | no/no | 1/0.3 |
| tw-005 | Dave Portnoy | “Who knows what's going on with the thespian… Inside the Actors Studio.…” | 2/1.35 | 1/0.04 | insult/factual_claim miss | no/no | 3/2.53 |
| tw-007 | Dave Portnoy | “I woke up to Bears QB @CALEBcsw spreading disinformation about me. I di…” | 3/2.26 | 1/0.37 | insult/insult | yes/yes | 2/1.04 |
| tw-008 | Craig Carton | “I don't quite understand why people immediately thought, 'Oh my God, it…” | 4/3.8 | 2/2.5 | insult/insult | yes/yes | 3/2.88 |
| tw-009 | Craig Carton | “He acted like his career ended… you're going to root for him because yo…” | 4/3.91 | 4/3.93 | insult/insult | yes/yes | 2/2.04 |
| tw-010 | Will Compton | “Now that we're on the other side of this news and Caleb Williams is oka…” | 2/1.22 | 2/1.46 | insult/insult | yes/yes | 1/1.02 |
| tw-012 | Will Compton | “We getting carted off the field for hamstrings now? We get in the ambul…” | 3/2.93 | 3/2.06 | insult/insult | yes/yes | 2/2.16 |
| tw-013 | Jason Whitlock | “America did not become the leader of the free world, and football did n…” | 3/3 | 4/3.98 | insult/insult | n/a | 0/0.06 |
| tw-014 | Robert Mathis | “Sorry @RGIII (all love lil bro) but I gotta chalk this one up in the so…” | 2/1.35 | 2/2.54 | insult/insult | n/a | 0/0.53 |
| tw-015 | Amani Toomer | “How are you gonna come in the huddle as a rookie. You got your nails pa…” | 3/2.89 | 3/3.16 | insult/insult | n/a | 0/0.22 |
| tw-016 | James Jones | “I'm on a different side… I am cool with crying, I've cried during footb…” | 2/1.63 | 0/0.19 | insinuation/insinuation | n/a | 3/2.99 |
| tw-017 | Boomer Esiason | “The level of entitlement is breathtaking — and it's no wonder why the c…” | 3/2.53 | 0/0.06 | insult/insult | n/a | 1/0.07 |
| tw-019 | Tyler Dunne (reporting anonymous Bears sources) | “Multiple Bears sources tell Go Long they've seen evidence that Williams…” | 1/0.36 | 0/0 | report/report | n/a | 0/0.01 |
| tw-021 | Merriam-Webster | “Just spelled 'conscientious' correctly on the first try.” | 2/1.03 | 0/0.22 | joke/joke | yes/yes | 1/0.35 |
| tw-022 | J.J. Watt | “A pulled hamstring while running full speed, not fun, hurts like hell……” | 1/0.46 | 0/0.07 | other/other | no/no | 0/0.01 |
| tw-023 | Dan "Big Cat" Katz | “There is a weird undercurrent of people online who have true Caleb Will…” | 0/0.61 | 0/0.02 | defense/defense | no/no | 0/0.01 |
| tw-024 | Maxx Crosby | “Everyone has an opinion about 'oh, why do you react this way to this ce…” | 0/0.15 | 0/0.01 | defense/defense | no/no | 0/0 |
| tw-025 | Kyle Long | “U made your post now shut the fuck up will … Nobody understands how muc…” | 1/1.05 | 0/0.08 | defense/defense | no/no | 0/0.01 |
| tw-026 | Robert Griffin III | “Watching Caleb Williams sobbing with his family after losing the game w…” | 0/0.01 | 0/0.01 | defense/defense | n/a | 0/0 |
Residual: parked tw-005 · scorecard.
/workspace/caleb-williams-fact-check/jev_stage_a/score_legends.md/workspace/caleb-williams-fact-check/jev_stage_a/frozen_posts.json/workspace/caleb-williams-fact-check/jev_stage_a/out/STAGE_A_FINAL_20260925.md/workspace/caleb-williams-fact-check/jev_stage_a/out/stage_a_run_20260925-115625_summary.json/workspace/caleb-williams-fact-check/jev_stage_a/out/stage_a_run_20260925-115625.jsonl/workspace/caleb-williams-fact-check/jev_stage_a/out/stage_a_lesson_draft.md/workspace/caleb-williams-fact-check/jev_stage_a/out/stage_a_explainer.htmlPaths are under /workspace/caleb-williams-fact-check/.
We built a tiny helper that spots rough public dunks about Caleb Williams.
It does not yell back on the internet by itself.
It finds a claim, wakes a human, and helps that human write a sourced pushback.
This page tells that story in plain words.
Next: Jev in one breath · Warm-up lesson: Stage A
Jev is a scoring model.
You feed it a short post.
It answers fixed questions.
It does not write a long chatty essay.
Think checklist, not debate club.
Same idea, five tone grades: what Jev is in the Stage A warm-up.
Next: Two labs
Stage A = warm-up
Stage A was practice.
We took 20 fixed posts.
Humans scored tone by hand (how loud, how gendered, joke vs insult, and so on).
Then Jev scored the same posts.
We rewrote the grading rules until the tone scores lined up.
Job of Stage A: learn the vibe of a dunk.
Not the full product yet.
Claim-loop = the real product test
The claim-loop asks the money question:
Tone rates and the five grades stay on the Stage A warm-up. They are a different test.
Next: How the claim-loop works
A rough post shows up.
Jev tries three things:
If yes, a human writes a short pushback with sources.
Nobody auto-posts. A person still hits send (or doesn’t).
Next: What we tested
When: Sep 26, 2026 (CT)
Cost: about $0.00123
Pack: 12 posts
about $0.00123 · typesafe/jev-1.13-20260917 on kind-of-post and worth-a-look · claim line via gpt-4o-mini workaround · run claim-loop-run-20260926-235043
Numbers for the three bars are on the honest scorecard. They come from SCORECARD_M2.md. We did not invent them.
Next: Honest scorecard
We set three bars. Fail any one → the loop is not “done.”
Numbers below come from SCORECARD_M2.md. We did not invent them.
| Bar | Need | Got | Result |
|---|---|---|---|
| Pull the claim on “should” posts | ≥ 7 of 8 | 2 of 8 | FAIL |
| Don’t false-alarm on controls | ≤ 1 of 4 | 0 of 4 | Pass |
| Human pushbacks hold up | ≥ 80% hold | 5 of 5 | Pass |
Bottom line: the full loop did not clear all bars.
The filter — don’t waste a person on a soft dunk — looks promising. The claim extractor, the piece that writes the claim down, is not ready as the spine. Stage A’s five grades are a different test. Those rates stay on the final scorecard.
Next: What worked · What failed
The filter is good.
On all 4 controls, Jev said “don’t wake anyone.”
Zero meme/insult-only junk in the human queue.
When a clear claim got through, humans finished strong.
5 drafts. All 5 held with receipts:
Drafts only. Person gate still on.
Next: What failed
Claim-pull broke.
We needed Jev to spot a checkable claim even when the tone is mean, then write that claim in one line.
It only cleared both steps on 2 of 8 “should” posts.
Often it called a fact “just an insult,” left the claim line blank, and still sometimes woke a human.
3 posts we thought deserved a fact-check never reached the queue (two Acho lines, one Esiason take).
Plain English: “Don’t waste humans” looks solvable. “Auto-extract the claim” is not ready to be the spine yet.
Next: Tiny example
Someone says crying on the cart is a “sacred rule” only for season-enders.
That is a checkable toughness rule, not just shade.
Humans already have receipts for that dunk.
Jev called it insult and said skip.
That’s the miss shape we must fix next.
Related warm-up walkthrough: Acho tw-001.
Next: What this is not
Not a pile-on machine.
Not auto-posting dunks back.
Not a claim that Stage A “proved the product.”
Stage A proved we can grade tone on a frozen set.
Claim-loop showed the product filter can stay clean, while claim extract still needs work.
Next: What happens next
Accuracy over heat. Especially on medical and anonymous-source stories.
Next: For builders
The five figures on this page fill the diagram slots: hero map, Jev in one breath, two labs, claim-loop flow, and scorecard bars.
Primary numbers are the locked M2 scorecard. Claim-pull FAIL 2/8. Control 0/4. Pushbacks 5/5.
jev_claim_loop/out/SCORECARD_M2.mdjev_claim_loop/out/pushbacks_m2.mdjev_claim_loop/out/misses_gold_yes_jev_no.mdjev_claim_loop/out/claim_loop_results_narrative.mdassets/hero-map.svg, assets/jev-one-breath.svg, assets/two-labs.svg, assets/claim-loop-flow.svg, assets/scorecard-bars.svgStage A paths stay on Files.