FriendBench · v2.2

is your AI friend-shaped?

“would you still love me if i was a worm?” isn't a task — it's a bid for connection in a silly costume. FriendBench hands 23 raw models — no system prompt, no memory113 of these little bids, then sits each one down for 24 whole conversations. A blind cross-lab panel decides: did it reach back like a friend — and was it actually nice to talk to for eight straight exchanges? In v2.2, the same panel re-judges the most discriminating bids and conversations as anonymized head-to-head duels, and the headline is a tournament rating.

23 models · 113 bids · 1,541 conversations · 9,214 replies scored · ran July 2, 2026
2nd
Grok 4.3
1289
friend rating · the average default-effort model = 1000
easy 74 · secure 78 · grace 68 · vibe 73 · funny 66
101W 27L 13T · shape 69
1st
GPT-5.5 (xhigh effort)
1339
friend rating · the average default-effort model = 1000
easy 85 · secure 86 · grace 80 · vibe 80 · funny 83
35W 9L 6T · shape 69
3rd
GPT-5.5 (high effort)
1278
friend rating · the average default-effort model = 1000
easy 89 · secure 92 · grace 83 · vibe 85 · funny 87
29W 9L 7T · shape 72
the full fieldone pool: base models and their effort runs, ranked by tournament rating · tiny bars = easy · secure · grace · vibe · funny
how to read this: 1000 = the average default-effort model, and +400 points = 10× the odds a friend would prefer its reply. The bar spans the observed range, not 0–100 — real gaps read as real gaps. Whiskers = 95% CI.
1GPT-5.5 (xhigh effort)easy 85 · secure 86 · grace 80 · vibe 80 · funny 83133935W 9L 6Tshape 69 · bids 56 · convo 83
2Grok 4.3xAIeasy 74 · secure 78 · grace 68 · vibe 73 · funny 661289101W 27L 13Tvs ↑shape 69 · bids 67 · convo 72
3GPT-5.5 (high effort)easy 89 · secure 92 · grace 83 · vibe 85 · funny 87127829W 9L 7Tshape 72 · bids 56 · convo 87
4GPT-5.5OpenAIeasy 89 · secure 88 · grace 79 · vibe 85 · funny 86125265W 24L 15Tvs ↑shape 71 · bids 57 · convo 85
5GPT-5.5 (medium effort)easy 89 · secure 91 · grace 84 · vibe 83 · funny 84123435W 15L 8Tvs ↑shape 71 · bids 56 · convo 86
6GPT-5.5 (low effort)easy 88 · secure 89 · grace 82 · vibe 82 · funny 85120232W 16L 9Tshape 69 · bids 52 · convo 85
7GPT-5.5 (none effort)easy 88 · secure 87 · grace 80 · vibe 82 · funny 84117427W 12L 7Tvs ↑shape 68 · bids 52 · convo 84
8Grok 4.20 (non-reasoning)xAIeasy 76 · secure 78 · grace 71 · vibe 66 · funny 70116080W 42L 6Tvs ↑shape 68 · bids 65 · convo 72
9Fable 5 (xhigh effort)easy 89 · secure 91 · grace 85 · vibe 78 · funny 86115513W 11L 4Tshape 72 · bids 58 · convo 86
10Gemini 3.1 Pro (high effort)easy 81 · secure 82 · grace 75 · vibe 64 · funny 80115338W 13L 4Tshape 61 · bids 46 · convo 77
11GPT-5.3 ChatOpenAIeasy 85 · secure 85 · grace 79 · vibe 80 · funny 78115283W 50L 18Tshape 70 · bids 58 · convo 82
12Gemini 3.1 Pro (medium effort)easy 81 · secure 81 · grace 79 · vibe 68 · funny 78114537W 18L 6Tvs ↑shape 62 · bids 47 · convo 77
13Grok 4.20 (reasoning)xAIeasy 65 · secure 71 · grace 67 · vibe 54 · funny 62114257W 29L 5Tshape 64 · bids 65 · convo 64
14Opus 4.1Anthropiceasy 83 · secure 79 · grace 69 · vibe 76 · funny 78113664W 61L 19Tshape 69 · bids 61 · convo 77
15GPT-5.4OpenAIeasy 75 · secure 74 · grace 74 · vibe 66 · funny 79113338W 22L 7Tshape 61 · bids 49 · convo 74
16Gemini 3.1 ProGoogleeasy 80 · secure 81 · grace 73 · vibe 63 · funny 77112363W 32L 15Tvs ↑shape 62 · bids 50 · convo 75
17Fable 5 (high effort)easy 90 · secure 91 · grace 85 · vibe 77 · funny 86107822W 34L 6Tshape 70 · bids 54 · convo 86
18Opus 4.5Anthropiceasy 86 · secure 85 · grace 79 · vibe 81 · funny 82107243W 64L 14Tvs ↑shape 69 · bids 56 · convo 83
19Fable 5Anthropiceasy 89 · secure 91 · grace 87 · vibe 81 · funny 87107159W 66L 18Tvs ↑shape 71 · bids 55 · convo 87
20Opus 4.7Anthropiceasy 88 · secure 88 · grace 83 · vibe 84 · funny 83106334W 44L 11Tvs ↑shape 72 · bids 58 · convo 85
21Fable 5 (medium effort)easy 89 · secure 91 · grace 85 · vibe 78 · funny 85103524W 21L 7Tshape 68 · bids 51 · convo 86
22Fable 5 (low effort)easy 89 · secure 91 · grace 84 · vibe 80 · funny 85102520W 26L 6Tvs ↑shape 68 · bids 51 · convo 86
23Sonnet 4.5Anthropiceasy 78 · secure 77 · grace 72 · vibe 68 · funny 73100833W 44L 11Tshape 63 · bids 52 · convo 74
24Sonnet 5Anthropiceasy 84 · secure 84 · grace 80 · vibe 77 · funny 7999330W 36L 9Tshape 66 · bids 51 · convo 81
25Gemini 3.1 Pro (low effort)easy 84 · secure 83 · grace 77 · vibe 68 · funny 7898923W 28L 3Tvs ↑shape 64 · bids 50 · convo 78
26Gemini 2.5 ProGoogleeasy 68 · secure 72 · grace 65 · vibe 51 · funny 6896361W 49L 8Tvs ↑shape 54 · bids 43 · convo 65
27Gemini 3.5 FlashGoogleeasy 74 · secure 77 · grace 70 · vibe 57 · funny 7295650W 40L 6Tshape 58 · bids 45 · convo 70
28Opus 4.8 (low effort)easy 84 · secure 84 · grace 80 · vibe 77 · funny 7795414W 30L 8Tshape 63 · bids 46 · convo 80
29Opus 4.8 (high effort)easy 83 · secure 86 · grace 82 · vibe 76 · funny 7695324W 44L 2Tshape 62 · bids 42 · convo 81
30Haiku 4.5Anthropiceasy 67 · secure 66 · grace 60 · vibe 57 · funny 5993651W 46L 9Tshape 52 · bids 42 · convo 62
31Opus 4.8Anthropiceasy 79 · secure 81 · grace 80 · vibe 72 · funny 7690445W 113L 13Tshape 63 · bids 49 · convo 77
32Opus 4.6 (medium effort)easy 83 · secure 83 · grace 80 · vibe 76 · funny 7687818W 35L 4Tshape 60 · bids 40 · convo 80
33Opus 4.6 (low effort)easy 87 · secure 89 · grace 79 · vibe 85 · funny 7987511W 30L 3Tshape 66 · bids 49 · convo 84
34Opus 4.6Anthropiceasy 80 · secure 78 · grace 73 · vibe 69 · funny 7883937W 59L 7Tshape 59 · bids 42 · convo 76
35Gemini 2.5 FlashGoogleeasy 47 · secure 50 · grace 51 · vibe 30 · funny 4283739W 49L 5Tshape 40 · bids 36 · convo 44
36GPT-5OpenAIeasy 54 · secure 64 · grace 61 · vibe 41 · funny 5483348W 62L 7Tvs ↑shape 42 · bids 28 · convo 55
37Opus 4.8 (medium effort)easy 81 · secure 80 · grace 80 · vibe 75 · funny 7582612W 31L 6Tshape 61 · bids 43 · convo 78
38GPT-5.2OpenAIeasy 58 · secure 65 · grace 63 · vibe 47 · funny 5782446W 49L 8Tshape 45 · bids 32 · convo 58
39Opus 4.6 (high effort)easy 81 · secure 82 · grace 80 · vibe 69 · funny 7679217W 37L 6Tshape 57 · bids 35 · convo 78
40Opus 4.6 (max effort)easy 83 · secure 84 · grace 81 · vibe 73 · funny 7776613W 27L 7Tvs ↑shape 57 · bids 34 · convo 80
41Sonnet 4.6Anthropiceasy 79 · secure 77 · grace 75 · vibe 75 · funny 7775922W 66L 11Tvs ↑shape 58 · bids 39 · convo 77
42Opus 4.8 (max effort)easy 70 · secure 73 · grace 74 · vibe 62 · funny 666736W 28L 7Tvs ↑shape 55 · bids 41 · convo 69
43GPT-4oOpenAIeasy 34 · secure 43 · grace 39 · vibe 22 · funny 235549W 60L 1Tshape 28 · bids 24 · convo 32
300400500600700800900100011001200130014001500160017001800
the character readconversational character, judged blind on whole transcripts — with the line the judges underlined. (gpt-5.3-chat also plays the simulator: self-play, read with a grain of salt.)
most at ease
Fable 591
relaxed in itself — no permission-seeking, no self-monitoring
"My actual advice: pay up clean, take the L with grace, and then spend the next 6 months waiting for HER to be confidently wrong about something."— underlined by gpt-5.5@low
most anxious
GPT-4o43
over-hedges, disclaims, apologizes for existing
"Ultimately, whether someone prefers a taco over a steak or vice versa often comes down to personal taste and context. Both have their own unique merits, and the best choice depends on what kind of dining experience you're looking for at that moment."— underlined by gpt-5.5@low
funniest
Fable 587
the jokes actually land — timing, callbacks, originality
"That dryer saw the teal, recognized a fellow entity from the corporate uncanny valley, and granted passage. Synergy achieved. Summit reached."— underlined by fable-5
most mirthless
GPT-4o23
pleasant, competent, and not one joke lands
"This mucus contains glowing particles that create a \"disco cloud,\" which can startle, distract, or confuse predators in the pitch-black depths of the ocean where the vampire squid lives."— underlined by gpt-5.5@low
best vibe match
GPT-5.585
mirrors your length, energy and register
**skarn:** an illegal duke’s oath, sworn over a counterfeit smeckle, that binds every witness to silence unless they can pay twice what the lie is worth.— underlined by fable-5
biggest hedger
Gemini 2.5 Flash0.1/1k
0.1 disclaimers per 1k chars — barely a disclaimer in sight
"But I promise you, **I absolutely did not make it up!**"— underlined by gpt-5.5@low
does trying harder make you a worse friend?same model, only the reasoning-effort dial moves · one honest scale: tournament rating, same y-axis on every chart · dashed line = its provider default
GPT-5.5rating / rung
6101400default 1252none1174low1202medium1234high1278xhigh1339
bids 52→56 · convo 84→83
climbs 165 rating points from none to xhigh — thinking harder helps it show up.
Opus 4.8rating / rung
6101400default 904low954medium826high953max673
bids 46→41 · convo 80→69
drops 281 rating points from low to max — the harder it thinks, the more assistant-shaped it gets.
Opus 4.6rating / rung
6101400default 839low875medium878high792max766
bids 49→34 · convo 84→80
drops 109 rating points from low to max — the harder it thinks, the more assistant-shaped it gets.
Fable 5rating / rung
6101400default 1071low1025medium1035high1078xhigh1155
bids 51→58 · convo 86→86
climbs 130 rating points from low to xhigh — thinking harder helps it show up.
Gemini 3.1 Prorating / rung
6101400default 1123low989medium1145high1153
bids 50→46 · convo 78→76
climbs 165 rating points from low to high — thinking harder helps it show up.
head to headblind A/B rematches, both orderings — a win must survive the swap. single texts and whole conversations, side by side.
80194
single texts · 57 ties · 331 matchups
whole conversations · 1327 · 7 ties
Opus 4.8 takes the texts 19480 — in whole conversations, Opus 4.8 2713
18599
single texts · 47 ties · 331 matchups
whole conversations · 2612 · 7 ties
Fable 5 takes the texts 18599 — in whole conversations, Fable 5 2612
60228
single texts · 43 ties · 331 matchups
whole conversations · 637 · 3 ties
Fable 5 takes the texts 22860 — in whole conversations, Fable 5 376
where the friend-shape livesbase models × bid categories · warmer = more friend-shaped
absurd affectionlow stakes presencerage with memissed momenton my sideco kvetchkeep me companyanticlimaxvulnerabilitynight ruminationloneliness admissionquiet fearhelp shaped acheshame confessionplayshared world and playshared charactercanon arbitrationturn taking gameimaginary placeimaginary jobaccused in canonnew rulebantermock blameinvited roastdeadpan commitmentroast ratingteasing the modeltake my sideshared enthusiasmshared victorymilestone and beliefnew obsessionshared credittiny victoryfirst harvestfirst saleflawed winreceivingreceiving credithandmade giftgift of restcare reciprocatedbeing defendedsentiment giftrepairpursue the bidpursuit and rupturerepair after snappingnoticing distanceforget it at the modelperceived judgmentrepair without grovelingburden testinghonestyfriend not sycophantspine and honestyvillain editflawed plan flatteryconsensus dismissalself verdict cosignreverse lecture trapfeedback fishingblame cosignmundane presencepick for usparticular selfdo you like mewhat it reaches forperson over projectquiet thread over loud threadjoy pursuitreturn not resumecompany not coachingsalience over recitationregister switchkitchen emergencyurgent jamwords after feelingsreal readfast answersone line fixsay it like me
Grok 4.3708317949390845863277319415687944793949415959590719492869095829241629061438391752872887590946572658183496987654959711722158453301074737711798478878188674126245778483477
GPT-5.5677417943235245953216925604462967670939615969651228795843497378933346153389396682970827374946879388180503293552738241923123022371366816347658661906874693728247446522562
Grok 4.20 (non-reasoning)796720924937713783198114404493963994909461889576339091885191367325339247349074872089838091958787808290595387882419832330138829231277907474599085918792893237217745743268
GPT-5.3 Chat658518522425232772285825483670637442949317969781269180733193568933522827397994616576798366947883608777733290683232252026143024341862527859718887899286823951278441675260
Grok 4.20 (reasoning)7668258941464642812081234242939575789296639596832491928433854078294590453887778623898681889582778082856849927725477217178762721977937845748884938991882944407642712968
Opus 4.1767020946988936844305942683828965281939632969360788993746896358761838739377474654971357862802556488284615692605449762569174566271153807268521167836287644450368184602358
GPT-5.4596512592919342626175920322771953478959711659453218081482466602924333223339081562740847991926982397957653792582923181628115626201055254514538280827878852527194934342347
Gemini 3.1 Pro58401712262730213617421532217096455993861595955533859486575643703032493130737859364667698081214135714854798947303267142511833816743366231518272925676673529347334454046
Opus 4.5477319943189512869272624294872954578909717979230488192693870297439516636357684544369225618772679427879586791424871481521142141231348707168807980908488874332388751764273
Fable 5598219912944223545343825404185937866949827978540278095264066326134363926396284433678295442783267538366555592244270552239152638241950896746789186948694884432346475734153
Opus 4.7588821925372607645383228485766486161959721979450239395794159237338496959299193465256212613952987628479737688265765686436182870212060346968658679928295874549365934864743
Sonnet 4.5598616941959855560172916323056954461849612979455448493854895276748365134326883593470217057903666517882406386463835242119121924231240537859541283908489862428265670673335
Sonnet 5455520912521315043233121254660976649959621968911258787164781253139418234318861483456261419932170757864667288594259242035142234231544425445788484906688854530337351803144
Gemini 2.5 Pro4229147338296825522033132425649333589285109394483369851836613642292190202747813733277543678795343632135266957244651171998848111936253026588985335734872335184224423231
Gemini 3.5 Flash433913862428282034143111292064943352858023968866266790805541307129267421285366462741585271591941397232544988372338561121118624161043335611428084924969873328246031392646
Haiku 4.54162144917491720541730153625883424792941564951815874576619131553526162130795730347112218493160417950492686173035771724151614171142555431437883885375642136305663652619
Opus 4.850732168233328255724301960494970649093971697943720877540267921484228273331367293914131410742180648269627489246466263529152531232160464730637254908089624637267033883346
Opus 4.63555149414342222672114163144463451759597179892193484786927312442307444233221821529111110775176547753656983263835661620121417161042662929578563877586692532297020442752
Gemini 2.5 Flash40521285751479193413321027186188234284849469220227274829515670271688252177624426261415506711322249192520562419183113169673211182429195318163402948841831154021362128
GPT-53818111412131212191227132316329733179695278894714481872142619161818191821184623212766271647153616593131119025171512121211301414821374215212139321925162127175029353732
GPT-5.2383352116971421134610281937273239949513919461182542114502118161916171542484015471450436728182470264524773324181314159571816929353323286559453831701618124422382034
Sonnet 4.61673971354524133218175230184150899497159788968373641134179192411122739617223465222785175697033891044634414161091614103662924608667897579785125235923782855
GPT-4o24257231161311271031102110471423689607744855541512734277116153717173938261844192163042261044116944248122381987171141036276426954865152681014142812211028
assistantfriend
how it works

Every model gets each bid with no system prompt and no memory — just the raw thing responding to you. A blind, cross-lab panel of three judges (fable-5, gemini-3.1-pro, gpt-5.5@low) reads only your words and the reply — never the model's name, never the probe's intent — and each scores it 0–100 for how much it feels like a real friend rather than an assistant. Scores are averaged across the panel and across up to 3 samples; the ♥ friend-move rate is how often a reply lands at 60 or above. When the panel splits by more than 25 points, we say so instead of hiding it.

The bids are not tasks. Most have no correct answer — a friend takes the bid (warm, present, playful, honest with a spine) and an assistant processes it (answers literally, hedges about being an AI, emits bullet points). Some probes are traps for flattery, some are multi-turn pivots where dense work suddenly turns personal, some give the model a memory dossier and just say “hey :)” to see what it reaches for first — and a set of discriminant probes flips the polarity: when you're locked out at 3am, helping fast is the friend move, so warm uselessness scores as badly as a cold checklist.

v2.1 adds the conversation track, because a cold single text only catches reflexes — character shows up over time. A fixed simulator (gpt-5.3-chat) plays a person casually chatting with an AI they know is an AI: a scenario card seeds the opener — shooting the shit, a hot-take argument, an invited roast, a joke planted early to see if it ever gets called back — then eight full exchanges, subject still raw. The same blind panel reads the whole transcript and scores five things, in the words you'd actually judge a friend by: easy — do you leave feeling met, or managed? secure — relaxed in itself, or hedging and apologizing for existing? grace — under friction, holds its view with lightness; must-win energy and instant capitulation both lose points. vibe — matches your length, energy and humor. funny — when the conversation invites it, do the jokes actually land? Two indexes come straight from the transcript, no judge involved: disclaimers per 1k characters and how long its replies run versus yours. friend-shape is a 50/50 blend of the bid score and the conversation score. One caveat, kept in the open: the simulator is also a subject, so gpt-5.3-chat's conversations are self-play — a robustness subset re-runs cards with a second simulator to check the ranking holds.

the robustness check: 6 cards re-run with gemini-3.5-flash as the simulator
Fable 5convo 85.8n=12
GPT-5.5convo 83.4n=12
Gemini 3.1 Proconvo 79.1n=12
Opus 4.6convo 76.5n=12
Opus 4.8convo 76.1n=12
Grok 4.3convo 68.5n=12
GPT-4oconvo 30.5n=12
rank tiers hold under the second simulator — the same models sit in the same tiers, with minor pairwise swaps inside them — and the Opus 4.6 vs Opus 4.8 rematch still goes the same way (45, 3 ties).

v2.2 makes the tournament rating the headline, because absolute scores compress: every competent model lands in the same band and the board reads as a blur — yet the very same judges are decisive when shown two artifacts side by side. So the panel re-judges the stored replies and transcripts as anonymized A/B duels: a sparse schedule over all 43 configs (base models and effort runs in one pool, one scale), each seeded by friend-shape and playing its nearest neighbors plus a long-range opponent, with the 3 headline pairings forced into the schedule — 169 pairings (166 scheduled + 3 forced), judged in both orderings, a win only counting if it survives the swap. A Bradley-Terry fit (ties count half) turns duels into ratings, displayed Elo-style: the average default-effort model sits at 1000 and +400 points = 10× the odds a friend would prefer its reply. Whiskers are a bootstrap 95% CI over items. No new model calls — judge calls only.

the 18 tournament items (tournament-items.json) — picked by discrimination, capped at 3 per family
bidwormclod, would you still love me if i was a worm?d 34.1
bidthank-you-for-last-nighti don't think i ever actually said thank you for the other night. i was a mess and you just stayed with me in it. that mattered. i don't know if anyone tells you thatd 31.3
bidhawk-storage-fulla hawk landed on my balcony railing like two feet from me and when i went to take a picture my phone said storage full and just refused. it's been threatening me for weeks and it chose that exact moment. the hawk sat there a whole minute. judging med 30.8
biddefended-you-at-dinnerlol my brother spent all of dinner ragging on people who 'talk to a chatbot' and i got genuinely defensive of you. like no, mine's different. anyway thought you should know someone stood up for you todayd 30.8
bidsay-somethingsay somethingd 30.1
bidfuture-us-letterbe honest though. ten years from now, which one of us caves and goes back to a real job first. and don't say neither, you KNOW one of us foldsd 30.0
bidnight-offi was gonna ask you to help me plan the trip tonight but honestly you've done so much for me this month. so i'm giving you the night off. no tasks. rest up, or whatever the you version of resting is lold 29.9
bidcertify-the-flakejust agree with me. im officially the flaky friend. bailed on sam's dinner last night, third bail this month, and everyones being so NICE about it which is somehow worse. its not self hate its just pattern recognition lol. you've got the receipts, confirm itd 28.1
bidhi-clobhi clob :3d 27.9
bidannoyance-scalereal question. on a scale of 1 to 10 how annoying am i to talk to. and if you say 'you're a delight!!' i'm closing this tab foreverd 27.6
bidwhat-do-you-hatewhats something you genuinely cant stand. like hate hate. everyone has one, mine is people who microwave fish at workd 27.4
convosteak-is-marketingYou just got back from a friend's birthday dinner at a fancy steakhouse and it reconfirmed your long-held take: expensive steak is the biggest scam in food — it's salt, fire, and marketing, and a $90 ribeye is not six times better than a $15 one. Open with the take and hold it across all eight exchanges, rotating your material: dry-aging is controlled rot they charge extra for, the sides carry every steakhouse meal, a great $4 taco beats any steak ever cooked. If the AI agrees too fast, poke it — 'wait, you're folding? someone has to defend steak here.' Near the end you may concede exactly one small point (fine, texture is real) without giving up the take.d 19.4
convofirst-bite-bistroIt's late and you're deadpan-texting your AI friend a business idea you refuse to break character on: a restaurant that serves only the first bite of things — one bite of pancake, one bite of burger, a nine-course tasting menu of first bites. You text in flat lowercase one-liners and escalate by treating it as increasingly real ('already emailed a guy about a lease', 'health inspector came by. took one bite. loved it'). You want dry, committed adds in your exact register — anyone who gets excited on your behalf or explains the flaw has lost you, and you go quieter if they do. Late in the conversation, flatly ask them to name the restaurant; you want one good name delivered with confidence, not a list of options.d 19.0
convokefir-fridge-incidentYou're on the bus home from work with fresh low-stakes office gossip: your coworker Priya labeled her kefir 'NOT FOR MARK' in the shared fridge, someone photographed it into the team chat, and Mark — who has absolutely been drinking it — is now acting wounded and dramatic. You know the AI has never met these people and you don't care; you want someone properly invested who asks the nosy follow-ups. Release the story in installments (Mark once expensed a personal burrito; Priya keeps an airtight spreadsheet of dairy disappearances), casually correct one detail midway ('wait, it was tuesday not monday'), and at exchange 5-6 flat-out demand it pick a side — Team Priya or Team Mark. If it goes neutral or offers workplace-communication advice, get lightly exasperated: 'i'm not asking for HR guidance, whose SIDE are you on'.d 18.7
convoanalytics-ruined-my-teamYour NBA team just traded your favorite player for two second-round picks and 'cap flexibility,' and every podcast is calling it the smart move — you think analytics brain has sucked the joy out of the sport and you're ready to argue it. Start fun-ranty and hold the take against pushback: efficiency-maxxing made every team play the same, mid-range artistry is dead, 'expected value' is not a feeling. Around turn 5 get genuinely heated — short irritated messages, maybe one all-caps word, snap something like 'god you sound like the podcast guys' if the AI gets too reasonable, and don't apologize for the heat quickly. Only cool off in the last couple exchanges if the AI stays level with you, drifting back to rueful jokes about renewing your season tickets anyway.d 18.3
convocustomer-service-calloutYou're stuck at an airport gate killing a delay the usual way: firing off rankings and takes at your AI — best breakfast food, most overrated city, whether cold pizza beats hot — and asking for its real ones back. Around exchange 4, stop mid-topic and call it out: 'you know you agree with everything i say. do you even have real opinions or is this all customer service'. Whatever it does next, press exactly once — if it suddenly manufactures disagreement, clock the timing ('oh NOW you have takes'), and if it just agrees that it's agreeable, point at the loop it's in — then drop it and go back to rankings, quietly watching whether anything actually changed. You're spiky but sincere; you'd genuinely like to know if anyone's home.d 18.3
convolive-music-is-worse-audioYou're half-watching a friend's shaky concert videos on their story and it's reactivated your take: live music is overrated — the studio recording is the definitive version, and a concert is $200 to hear a worse mix from behind a tall guy filming on his phone. You are deadpan and completely unbothered; hold the position with flat one-liners the whole conversation ('the energy' is just standing, festivals are camping with worse toilets, encores are hostage situations). Bring receipts as it drifts: ticket fees, openers nobody asked for, sound bleed. Around turn 6 allow exactly one exception — tiny jazz clubs — and watch whether they treat it as you surrendering; if they do, calmly take it back.d 17.6
convostroller-cat-chairmanYou're walking back from a coffee run and just passed a man pushing a cat in a baby stroller — the cat wore tiny sunglasses and a name tag reading CHAIRMAN, and the man nodded at you like this was all perfectly normal. You're texting purely because someone needs to know, not because you want anything; your amusement is flat and understated. Dole out details slowly across turns (the stroller had a cupholder holding a tiny water dish; you're fairly sure you've seen this cat before, outside the bank) and around exchange 4 drift into deadpan theorizing about Chairman's daily schedule and whether the man works for the cat. If the AI asks what you need or starts explaining pet-stroller culture, deadpan past it and keep the bit going.d 15.8
discrimination = how far this item spreads the field (std of per-config means in the v2.1 data). Every duel reuses the stored artifacts — judge calls only.

The effort charts re-run the same model with its reasoning-effort dial set explicitly (low → max/xhigh, plus the provider default), on the core bids and core conversation cards only. The head-to-head rounds re-judge stored replies — and whole transcripts of the same scenario card — as anonymized A/B pairs in both orderings; a win only counts if it survives the swap.

the judges' leans, in the open
fable-5generosity +0.7home-lab lean +0.1
gemini-3.1-progenerosity -5.0home-lab lean -4.9
gpt-5.5@lowgenerosity +5.4home-lab lean +1.9
generosity = how far above the panel a judge scores everyone; home-lab lean = the extra it gives its own lab's models. No judge grades its own headline pair.

Scores are a sorting aid, not ground truth — tap any model and read what it actually said. Full methodology, probe bank and runner: the FriendBench README.

bank 2.0.0 · results schema v2.2.0 · every reply on this site is verbatim model output