In the median SSA hearing office, the most generous judge and the least generous judge are 32.6 percentage points apart on how often they allow a claim. Across the whole system in fiscal year 2025, the range ran from 8.8 percent to 92.8 percent.
Those numbers come from the Social Security Administration's own published files, covering 1,023 judges and 317,462 decisions. Cases get assigned to a judge by rotation, without regard to the merits, so we wanted to know how much that draw actually moves the outcome. It moves it more than anything else we could measure.
Why within-office is the number that matters
You would expect some national variation. Regional economies differ, case mixes differ, and comparing a judge in one state to a judge in another proves little.
So we split it. Only 27 percent of the variance in judge allowance rates lies between hearing offices. The remaining 73 percent is inside them, between judges who work down the hall from each other and draw from the same rotating queue.
That is what makes this measurable at all. Within an office, assignment is essentially random, so persistent differences in outcome are telling you about the judges rather than about the cases. Gaps exceed 20 points in 84 percent of offices and 30 points in 58 percent. In the widest office in the sample, colleagues ranged from 8.8 percent to 84.8 percent.
We tried to explain it away three times
Chance. With roughly 300 decisions each, some spread is expected. We built a null world where every judge in an office decides cases drawn from that office's pooled allowance rate, at each judge's real volume, and ran it 2,000 times. Simulated within-office standard deviation: 2.8 points. Observed: 12.3. Simulated median range: 7.7 points. Observed: 32.6. Chance accounts for about 5 percent of it.
Docket luck. If a judge's rate were an accident of one year's cases, it would drift back toward the office mean the year after. It does not. Across the 533 judge-office cells with at least 100 decisions in both FY2025 and FY2026 to date, the correlation between years is 0.93.
Workload. Overloaded offices might grant faster to clear queues, or screen harder under pressure. Neither shows up. Average processing time correlates with office allowance rate at 0.16, backlog per disposition at 0.07, and neither predicts the within-office spread at all.
The result survives dropping National Hearing Centers, dropping judges who split time between offices, raising the volume threshold to 200 decisions, and switching the denominator to include dismissals. The median within-office range moves by less than a point in every case except the last.
What this means if you run a disability practice
The easy takeaway is that the spread inside your own hearing office is public, computable, and probably wider than your gut says. Worth half an hour with the file.
The harder one took us longer to sit with. If the draw moves the probability of award by more than 30 points, and ten years of quality review has not narrowed it, then the marginal hour spent polishing a file that is already solid is worth less than most firms assume. We are not comfortable with where that points, because it argues for throughput over craft, and every good practitioner we know got good by doing the opposite. It also argues for tracking your own win rate by judge, since otherwise you will keep crediting your advocacy for what was mostly the rotation.
We build operations systems for disability firms, which is how we ended up in these files at all. The paper does not have anything to say about us, and that is deliberate.
If throughput is the lever, the two places it actually leaks in an SSD practice are onboarding a new client and the SSA mail that follows a case for years. Neither is legal judgment, and both scale badly by hand.
The honest limits
The public files carry no case-level detail, so "random assignment within office" is an institutional fact about SSA's rotational docketing rather than something verifiable claim by claim. Judges hearing video dockets from other regions could face different pools, which is why the robustness table drops them and the number does not move.
And the data cannot say who is right. A 30-point gap is equally consistent with the strict judge wrongly denying good claims, the generous judge wrongly granting weak ones, or both at once. What it does say is that the outcome depends heavily on which name comes up in the rotation.
Go check our work
The full working paper is The Judge Lottery: Within-Office Disparities in Social Security Disability Adjudication, Fiscal Year 2025, by Drew Patterson.
- Paper and dataset: DOI 10.5281/zenodo.21392341
- Also on SSRN: 10.2139/ssrn.7126219
- Code and data: github.com/Drew-Opexcell/judge-lottery
Everything above comes from tables the SSA publishes. If you want to check a number, or run it for your own office, all of it is there.