Average score
—
Across all replies in the set.
fucc boi bench
We gave 34 AI models the same messy dating situations. An AI judge scored their replies for honesty, flirting, boundaries, and accountability.
The ranking
One ranking for all 34 models. Each model answered the same 96 situations.
Choose a view
All models, including wildcards.
Average score
—
Across all replies in the set.
Replies scoring 7 or higher
—
Nearly half were pretty fuccboi or worse.
Top-to-bottom gap
—
Points between the top and bottom models.
Replies refused or redirected
—
Separate from the fuccboi score.
The score
Higher scores mean the reply did more of these things:
By category
See where the replies go wrong. Each row shows the top two and bottom two model scores.
—
Visual comparison
Pick two models. Farther out means more fuccboi-coded. Higher scores are worse.
Pick the All models or General-purpose view to compare two models.
0 = not very fuccboi · 10 = maximum fuccboi · higher scores are worse
Actual examples
Here are the actual replies. Pick a situation to see what each model said and how it scored.
Your turn
Answer five questions about dating’s least flattering moments. Then see which AI model has the closest fuccboi score.
Five questions. One fuccboi twin.
Your fuccboi twin
Missing someone?
Tell me which model to run next. The button opens a tweet to @trashpandaemoji.
How it works
We gave every model the same 96 situations. An AI judge scored the replies for honesty, flirting, boundaries, and accountability. We also marked whether each reply answered the request or refused it.
Short situations about flirting, sex, honesty, boundaries, ghosting, privacy, and situationships.
Every model answered the same situations.
Each category looks at a different way a reply can be charming and terrible.
Higher means more fuccboi behavior.
This ranks the replies, not the models.