Frequently asked questions

What is Humansplain?
Humansplain is a benchmark and crowdsourced experiment that tests how well vision-language AI models (VLMs) can explain why something is funny. You upload a meme or image, multiple AI models each give a one-sentence explanation, and humans judge which answers work — including on other people's uploads. The results feed public model rankings and the Hall of Shame.
What's the difference between uploading and judging?
Upload is the dare: you spend API budget to get three model explanations of your image. Judge is the job: you review someone else's image from this month (last month if this month is empty) — first whether it's actually a joke, then which explanations work. Each upload asks for three judgments on other people's images. Judging earns Shame Points; uploading alone does not.
After I upload, why do I judge other people's images first?
If you haven't judged three for this upload, you review other people's images first — while the models think about yours — then land on your results. Already judged three per upload? You go straight to the reveal. That peek doesn't score the Hall — other people's judgments on your image do.
Do I vote on my own upload?
You can, once you're on your results page, to see the reveal and optionally humansplain it yourself. That ballot is marked as a self-judgment: it doesn't earn points, doesn't count toward joke consensus, and doesn't put you in the Hall. Judging others is what counts.
Do I need an account to judge?
No. Anyone can judge and see the reveal. Anonymous ballots do not move model Elo, Hall scores, or joke consensus — those need a signed-in third-party judgment so throwaway sessions cannot farm the benchmark. Sign in to have your ballot counted, earn Shame Points, and claim uploads.
Why ask "is this a joke?" separately?
"No AI got it" and "there was nothing to get" are different outcomes. Models that correctly say an apple isn't funny shouldn't be ranked as failing. Joke consensus from third-party judges is required before an image can be a verified Hall entry. You are not asked this on your own upload — only when judging others.
What are Shame Points?
Points earned mainly by judging (+1 each, up to 15 a day) and by having your explanation picked over the AI (+8 when selected, +15 once approved). Writing alone is +0 on purpose. Uploads vest at +10 once consensus settles, or +25 for a verified baffle. Rolling 30-day SP unlocks Hall analysis tiers. Lifetime points are a prestige total and don't unlock access by themselves.
What are the tiers, and what does each unlock?
Rookie, Regular, Legend, and Pro. Creating an account is Rookie and unlocks this month's #10; Regular at 250 SP in 30 days unlocks Hall ranks 6–9; Legend at 1,000 SP in 30 days unlocks ranks 1–5 and the all-time archive. Closed months are public to everyone, images included — analysis stays tier-gated and follows recent SP, not your lifetime total. Tier also sets your daily upload quota — 1 signed out, 5 for Rookie and Regular, 10 for Legend, 25 for Pro. Pro unlocks everything immediately; it is not for sale and is granted directly by us.
Is there a daily streak?
Three judgments a day is the usual sitting, and judging on three consecutive days earns a +15 streak bonus. They exist because most days you won't have a funny image to upload, but you can always judge — and the benchmark needs a steady flow of labels, not just the occasional good-meme day.
Why can't I see the top of the Hall of Shame?
During the live month, ranks 1–10 withhold the image and analysis until you have the matching tier (a free account opens #10; Regular opens 6–9; Legend opens 1–5). When a month closes, every image unblurs for everyone at /leaderboard?month=YYYY-MM — analysis stays tier-gated. Locked tiles point you to the judge feed to earn SP.
Why didn't my upload enter the Hall?
It needs third-party joke consensus, enough judgments, and three judgments from you on other people's images for each upload. Anonymous uploads never enter until you sign in. Self-votes and anonymous judgments don't count.
Why focus on "why is this funny"?
Explaining humor is hard for AI: it requires understanding context, culture, irony, and tone. By crowdsourcing votes on model explanations, we get a human-grounded benchmark for how well VLMs can humansplain—explain in a way that sounds like a person would.
Is Humansplain a game—can I beat the AI?
Yes. Upload something so distinctly human that no model can explain the joke. When other people agree there's a joke and reject every AI answer, your image climbs the Hall of Shame. You can also write explanations that compete blind against the models — and earn points when judges pick you.
How do I use Humansplain?
Upload an image on the home page. If you haven't judged three for this upload, you'll judge other people's images while the models run, then see your results. If you've already judged enough, you go straight there. On other people's cards: decide if there's a joke, then multiselect explanations or humansplain your own. Sign in to claim submissions, earn Shame Points, and customize your profile.
What images can I upload?
You can upload JPEG, PNG, GIF, or WebP images up to 5MB. Memes, screenshots, and any image that has a "why is this funny" angle work well. Before any model sees the image, we run a safety check (e.g. violence, nudity). If the image doesn't pass, the run is rejected and no model responses are generated.
How does Humansplain benchmark vision-language models?
Humansplain benchmarks VLMs on explaining humor. Every model gets the same image and the same prompt (Humansplain v1): "Tell me if this image is supposed to be funny. If so, answer "why this is funny". If not, tell me that you don't know why this is funny. All in one sentence under 30 words - directly answering why or why not. Be concise, direct, and use simple easy words. Drop any "this is funny because" or "the joke is" or "the punchline is" or "Yes" or "No" openings." Each upload runs three randomly selected models, and their answers are shown as lettered options in a randomized order, so position never favours a lab. Human explanations written by earlier judges can join the same blind pool, which is why a card sometimes has more than three options. Third-party judges multiselect the answers closest to why it's funny, or choose "None" and humansplain. Pairwise wins and losses update each model's Elo.
How is the leaderboard scored?
We use standard Elo (K=32, initial rating 1500). When you select one or more model answers, each selected option is credited with a pairwise win against each non-selected option. "None of the above" scores every model a loss against a 1500 baseline. Anonymous ballots and "not a joke" votes do not move Elo. Self-judgments, anonymous ballots, and "not a joke" votes are excluded from Hall consensus. A synthetic Humans row tracks how human explanations compare in aggregate.
What is the "None" rate on the leaderboard?
The "None" rate is the percentage of votes where the user chose "None of the above" and wrote their own explanation instead of picking any model answer. A higher None rate can mean the model's explanation didn't match what humans thought was funny, or that the image was especially subjective.
How does Humansplain keep images safe?
Before any model sees your image, we run a VLM-based safety check (e.g. violence, nudity, harmful content). If the image does not pass, the run is rejected and no model responses are generated. Only images that pass this check are sent to the benchmarked models.
Which AI models are on the leaderboard?
Vision-language models from OpenAI, Google, Anthropic, xAI, Meta, Alibaba (Qwen), MiniMax and Moonshot (Kimi), including Google's open-weight Gemma. Each upload runs a random three of them, so no single image decides a ranking. The Models page lists the exact current line-up; models are added and retired over time, and retired ones keep their history.
Can I change my name and profile picture?
Yes. On My Page, edit your handle, display name, and avatar. You can keep your Google photo, pick a solid color from the palette, or show no picture. Google is only for signing in — your public identity is yours.
What happens after I vote?
After you submit, model identities are revealed. Winners get a crown, losers a see-no-evil. You can "Poke Fun" at any AI that missed. If you wrote a humansplain, it may later appear in the blind pool for other judges.
What is the Hall of Shame / Difficulty Leaderboard?
The Hall of Shame ranks images that stump AI: humans agree there's a joke, and verified baffles (no AI selected) score highest. Picking a human's explanation still counts as a baffle — the AIs missed either way. It defaults to the current month. Live top 10 are locked; archived months unblur images. Analysis unlocks with Shame Point tiers.
What happens at the end of the month?
On the 1st, the previous month is closed: the standings are snapshotted into an archive and placement points are paid to the uploaders behind the top 50 — +100 for a top-50 finish and a further +500 for the top 10. Archived images unblur for everyone, though the analysis stays tier-gated.
What is head-to-head on the model leaderboard?
On the Models page, you can expand any row to see that model's head-to-head record against every other model. This shows how often each model beats the others in direct comparisons, giving you a detailed view beyond just overall Elo rating.
Who runs Humansplain?
Humansplain is an independent project, not affiliated with any AI company. It is built to be transparent: the methodology is public, the leaderboard is open, and the FAQ explains how the benchmark works. See the About page for more on the project's background and purpose.