LeetCode's Difficulty Labels Are Stuck in 2015 — Here's a Better Way to Score Problems
LeetCode's Easy/Medium/Hard tiers are static thresholds on aggregate stats, which misrepresents difficulty across skill groups. This post proposes a hybrid score combining solve time, success rate, and problem length, with a worked example and validation steps.
11 Jan 2026, 07:20 UTC

Pick a LeetCode Medium at random and you'll likely get one of two things: a ten-minute warm-up or a disguised Hard. The Easy/Medium/Hard label tells you almost nothing about whether you will find it difficult, because the label isn't really about you at all. This post argues that static tiers are a weak proxy for difficulty, and shows how to build a data-driven alternative you can adapt for your own practice tracker, team leaderboard, or study plan.
How LeetCode Assigns Difficulty Today
LeetCode's three tiers are effectively static thresholds over aggregate statistics — average solve time and success rate. The label doesn't know whether you're a competitive programmer or a first-timer, and it doesn't change as the community's skill distribution shifts. A graph problem labeled Hard in 2016 might be routine today because every interview-prep guide covers it, while an obscure bit-manipulation problem stays Hard mostly because few people attempt it.
The blind spots are consistent: problem length is ignored, language overhead is ignored (a Rust borrow-checking fight is not the same as a Python one-liner), and community feedback — upvotes, downvotes, discussion sentiment — plays no role. Niche topics like concurrency get distorted because the population of attempters is self-selected.
A Hybrid Difficulty Score
A more balanced approach combines three signals:
- Average solve time — how long accepted submissions take, per problem.
- Success rate — accepted submissions divided by total attempts.
- Problem length — statement size as a proxy for specification complexity.
Normalize each signal to [0, 1] across your problem set, then combine:
score = 0.4 * norm_solve_time + 0.4 * (1 - norm_success_rate) + 0.2 * norm_lengthThe weights are tunable. Solve time and success rate carry the most signal; length is a tiebreaker for problems with sparse data.
A Worked Example
Suppose you have per-problem stats in a CSV: slug,avg_solve_time,success_rate,length. In Python:
import pandas as pd
df = pd.read_csv("problems.csv")
# Min-max normalize each column
df["n_time"] = (df.avg_solve_time - df.avg_solve_time.min()) / (df.avg_solve_time.max() - df.avg_solve_time.min())
df["n_fail"] = 1 - (df.success_rate - df.success_rate.min()) / (df.success_rate.max() - df.success_rate.min())
df["n_len"] = (df.length - df.length.min()) / (df.length.max() - df.length.min())
df["difficulty"] = 0.4 * df.n_time + 0.4 * df.n_fail + 0.2 * df.n_len
# Map to tiers at the 33rd and 66th percentiles
q33, q66 = df.difficulty.quantile([0.33, 0.66])
df["tier"] = pd.cut(df.difficulty, [-0.01, q33, q66, 1.0], labels=["Easy", "Medium", "Hard"])Run this anywhere Python and pandas are available; no special permissions needed. Expected check: the resulting tier counts should differ from LeetCode's official labels for a meaningful fraction of problems — that's the point, not a bug.
Making It Dynamic
Recompute the score as a moving average over the last N attempts rather than all-time stats. This lets difficulty adapt as the community's skill distribution drifts. It also enables targeted recommendations: nudge users toward problems slightly above their current level instead of a static tier.
Trade-offs and Limitations
LeetCode's internal scoring is proprietary, so the hybrid metric is an approximation, not a reproduction. Solve-time data may be skewed by language mix — a problem mostly attempted in C++ will look slower than one attempted mostly in Python. And dynamic labels introduce churn: users accustomed to "Medium" may be confused when a problem's difficulty changes week to week. A practical mitigation is to display a score with a confidence band rather than a hard tier.
How to Check Your Work
Collect 500+ problems, compute both the official tier and the hybrid score, and compare distributions with box plots of solve time per tier. If the hybrid metric redistributes problems sensibly — e.g., long, low-success-rate problems move from Medium to Hard — the score is doing its job. A small usability study where participants rate perceived difficulty against the score is the strongest validation.
The actionable takeaway: don't trust a tier label as a measure of your difficulty. Compute your own score from your own attempt history, and let it update as you improve.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.