← Back to Blog
Game Design

Difficulty Curves 101: How Easy, Medium, and Hard Trivia Categories Are Built

2026-07-05·6 min read

Ask a room of trivia writers to sort fifty questions into easy, medium, and hard, and you will get fifty different opinions and at least three arguments. Everyone agrees difficulty is real — some questions clearly land harder than others — but almost nobody agrees on what is actually being measured when they call a question hard. That disagreement is the whole problem, and understanding it is the first step toward building a difficulty curve that actually works instead of one that just feels arbitrary to the people playing it.

Difficulty is not the same as obscurity

The most common mistake in sorting trivia questions is treating difficulty as a direct measure of how obscure the fact is. Under this logic, a question about a country's capital is easy, and a question about a minor supporting character in a decades-old film is hard, simply because fewer people could name it cold. That instinct is not wrong exactly, but it is incomplete, and leaning on it too heavily produces a curve that punishes players for not sharing the question writer's specific interests rather than testing anything resembling general knowledge.

A better way to think about difficulty is as a prediction about how many people in a broad, mixed audience could answer correctly, not how deep the fact sits in some imagined hierarchy of knowledge. A question about a mid-tier player from a regional sports league might be genuinely obscure to most people and still land as easy in a room full of that sport's fans, while a seemingly basic question can trip up a surprising number of people if it depends on a detail nobody actually rehearses, like the order of events in a story everyone technically knows. Difficulty, in other words, is a property of the audience as much as it is a property of the fact, and any curve that ignores this ends up miscalibrated for whoever is actually playing.

The false difficulty trap

There is a second, sneakier failure mode that has nothing to do with how obscure a fact is: questions that feel hard not because the knowledge is rare, but because the question itself is poorly built. Ambiguous wording, double negatives, answers that depend on an unstated assumption, or dates that could plausibly round either way all inflate a question's apparent difficulty without adding anything meaningful to it. A player who gets this kind of question wrong has not been tested on their knowledge; they have been tested on their ability to guess what the question writer meant.

This distinction matters enormously in practice, because false difficulty and real difficulty feel identical from the outside — both produce a low correct-answer rate — but they produce completely different experiences for the people playing. Real difficulty, even when a player misses it, tends to land as fair. There is a small flicker of respect for the question, an "oh, that's a good one" reaction, even in defeat. False difficulty lands as frustrating, because the miss feels arbitrary rather than earned, and a curve built on a pile of ambiguous questions will feel harder and worse than one built on genuinely obscure but cleanly written ones, even if the raw correct-answer percentages come out identical.

Building the actual curve

Once you can reliably tell real difficulty from false difficulty, the next problem is arranging questions so the curve does something useful across a full round rather than just existing as three labels on individual questions. A curve that front-loads its hardest material punishes players before they have found any rhythm, and a surprising number of people quietly disengage from a game in the first sixty seconds if the opening questions make them feel outmatched. Starting easy is not about being condescending; it is about giving everyone, regardless of how much they end up scoring, an early moment of getting something right, which buys enough goodwill to keep them engaged through the harder stretch that follows.

The middle of a well-built curve is where most of the actual game happens, and it deserves more attention than it usually gets. Medium questions are the hardest category to write well precisely because they have no obvious identity of their own — they are defined entirely by not being easy and not being hard, which makes them easy to fill lazily with whatever didn't clearly sort into the other two piles. A strong medium question usually has one identifiable hook, a detail that a reasonably engaged but not expert player might recall if they think about it for a few seconds, rather than either an instant recall or a total blank. That brief moment of active thinking, rather than instant recognition or instant surrender, is what separates a good medium question from a filler one.

Hard questions, meanwhile, earn their place by rewarding real depth rather than just rewarding rarity for its own sake. The best hard questions in a curve tend to be ones where even players who miss them learn something interesting in the process, and where the handful of people who do get it right feel genuinely, visibly pleased with themselves. That reaction is worth designing for deliberately, because it is often the moment a casual player becomes a repeat player — the game gave them one moment to shine, and people remember that far longer than they remember losing.

Categories complicate everything

Difficulty curves get considerably harder to manage once you introduce categories, because a category's internal difficulty spread rarely matches another category's spread in any clean way. A history category can range from schoolbook-basic to genuinely specialist within just a few questions, while a category built around a single narrow pop-culture franchise might not have much room between its easiest and hardest questions at all, simply because the whole subject sits at a similar altitude of familiarity. Treating every category as if it can supply an identical easy-medium-hard spread produces some categories that feel artificially padded at the easy end and others that feel artificially inflated at the hard end just to hit a quota.

The more honest approach is to let category composition vary and compensate for it in how categories are sequenced across a full game, rather than forcing every single category to internally mirror the same difficulty distribution. A category that runs slightly harder overall can be balanced by placing it after one that runs easier, so the felt difficulty across the whole session stays roughly even even when no individual category is perfectly graded on its own.

Calibration is never really finished

None of this works as a one-time design exercise, because difficulty is ultimately an empirical claim about how real people respond, not a property you can fully determine by reasoning about a question in isolation. A question that looks obviously medium on paper can turn out to play easy once actual people see it, or the reverse, and the only real way to find out is to watch it get played and see what the correct-answer rate actually looks like. Question sets that never get revisited after their first draft tend to drift out of calibration over time anyway, since what counts as common knowledge shifts as the culture around it shifts, and a question that was safely easy a few years ago can quietly become a stumper as its reference point fades from daily relevance.

That is really the underlying truth behind any working difficulty curve: it is less a fixed structure you build once and more a live hypothesis you keep testing against an actual audience. The three labels — easy, medium, hard — are a useful shorthand, but the real craft is in the constant, slightly humbling process of checking whether the questions are doing what you thought they would, and adjusting the ones that quietly aren't.

← Back to Blog