
Systems & Institutions
06 July 2026
16 September 2026
Sean William Hammond
Judges Got a Sentencing Database – Measuring Judges Became a Crime
Judges Got a Sentencing Database – Measuring Judges Became a Crime
In 2017 I proposed two things. One became a federal tool that judges use tens of thousands of times a year. The other became a crime in France, and in the United States it was never even refused. It was absorbed into a letter from 1988 that nobody has had to defend.
Both halves used the same data and the same arithmetic. The only difference was who got measured.
That difference is the most reliable thing I know about how institutions adopt new technology. The version that gives decision-makers better information gets built, because the people who approve it benefit from it. The version that shows the public how those decision-makers actually perform, doesn't. That would require those same people to consent to being watched. And consent to evaluate one's performance is not usually forthcoming.
Watch which half gets built. It tells you what the institution is optimizing for, no matter what the press release says.
Courts are the cleanest case I know, because I have been following this one for nine years.
The proposal
I was finishing a philosophy degree and reading the Herald-Tribune investigation into Florida sentencing. Their finding was blunt. Trial judges across the state gave black defendants longer sentences than white defendants — sometimes double — for the same crimes under identical circumstances.
My argument was that we already had the instrument and were refusing to use it.
Fourteen years earlier, Billy Beane had taken the poorest team in baseball to the playoffs by noticing that a century of expert scouting was systematically wrong in ways nobody had bothered to count. Moneyball made him famous. What made him right was simpler than the legend: he counted things the experts had decided did not need counting, and the counting won.
Every professional sport now runs on that principle. We built cameras to settle arguments about a game. We have almost none of that apparatus for prison.
So I proposed two things.
Give judges a standardized baseline showing what similar cases actually receive, so a departure from the norm is visible as a departure.
Then measure the judges. Track how often each one goes above or below that baseline, for whom, and by how much.
The courtroom, I wrote then, is not trying to predict the future. Quantifying the past is where this kind of analysis works. I did not know how much weight that sentence would carry.
The half that got built
In September 2021 the United States Sentencing Commission launched Judiciary Sentencing INformation. A judge enters the guideline, the offense level, and the criminal history category, and gets the national average and median sentence imposed on similarly situated defendants over the previous five years.
Public. Free. In fiscal year 2025 the Commission reported that JSIN was accessed about 4,800 times a month by nearly 17,000 users.
And in March 2023 the Federal Judicial Center began a two-year randomized pilot, appending JSIN information to presentence investigation reports in a set of districts, with the rest as controls. They formally tested whether showing a judge the baseline changes what the judge does. I have not seen the results published. The fact of the trial is the point: an institution decided the question was worth answering with a control group.
That is half the 2017 proposal, implemented federally.
Its critics are right about its flaws. It includes mandatory minimums, which are not exercises of discretion. It excludes cooperating witnesses, which distorts comparisons for defendants who wanted to cooperate and had nothing to trade. It reports national figures only, washing out the regional variation that often is the disparity.
All fixable. What matters is the direction. Judges should see what other judges do. Someone built that, published it, and put it through a controlled trial.
I did not know it existed until I went looking this year.
The half that became a felony
In 2019 France passed Article 33 of its justice reform law. In substance: the identity data of magistrates and members of the judiciary may not be reused with the purpose or effect of evaluating, analyzing, comparing, or predicting their professional practices.
The penalty tracks the personal-data provisions of the penal code. The legal press reported it as five years in prison. It applies to individuals, researchers, journalists, and companies. It was the first law of its kind anywhere.
It arrived inside a reform that otherwise expanded public access to court decisions. France made the rulings more available and made analyzing the people who issued them a crime, in the same document.
The stated justification includes preventing litigants from shopping for favorable judges. That concern is real, and it answers itself. Forum-shopping only works because sentencing is inconsistent. Publish the inconsistency and the pressure runs toward fixing it. A judge who is reliably lenient and a judge who is reliably severe are the same failure in different clothes. The remedy is identical: show the pattern, let it be explained or corrected.
Reporting at the time also indicated that judges pushed for the provision after publicly available information suggested certain magistrates were routinely rejecting asylum applications. If that reporting is right, that is not a privacy interest. That is a finding. The response to the finding was to criminalize finding it.
Article 33 bans publication. It does not stop a judiciary from auditing itself internally — which is worse, not better. An institution that measures itself and publishes nothing is not accountability. It is self-assured. Accountability requires somebody on the outside who can check.
Every profession that has resisted measurement used the same language: the work is too nuanced, the numbers cannot capture it, outsiders will misread the data. Medicine said it. Psychotherapy said it. Baseball scouts said it for years. In every case the resistance was sincere, the measurement came anyway, and the practitioners who adopted it got better at the job.
A bad scouting report costs someone a roster spot. A bad sentence costs someone a decade behind bars.
The half that got quietly withheld
France at least passed a law. The United States handled it with a letter.
Under 28 U.S.C. § 994(w), the chief judge of every federal district must send the Sentencing Commission a complete report on every criminal sentence within thirty days: the judgment, the written statement of reasons, any plea agreement, the charging document, and the presentence report. Every case. The Commission codes it. In fiscal year 2025 that meant 66,662 individual cases.
That database contains everything needed to evaluate a judge. Which judge, what offense, what guideline range, the defendant's demographics, the sentence imposed, and the judge's own written reasons for any departure.
The public data files are stripped of identifiers.
The Commission explains why on its own site. Consistent with a memorandum of understanding with the Administrative Office of the U.S. Courts, it does not release information identifying an individual defendant or any other person identified in the sentencing information.
The agreement is a letter dated June 22, 1988, countersigned July 21, 1988. It exists to protect the confidentiality of presentence reports, which are genuinely sensitive documents. Neither FOIA nor the Privacy Act applies to the federal judiciary. Judges are not its subject. They fall under "other person identified," a phrase broad enough to include the person who imposed the sentence.
In thirty-eight years, nobody has revisited it.
A deliberate decision can be reversed by whoever made it. This is nobody's decision. It has no author sitting in the room, no defender, and no process for undoing it — and it accomplishes exactly what France needed a criminal statute for.
The federal defenders' Sentencing Resource Counsel has noted that the Department of Justice and Congress receive access to Commission data that the public does not. If that is right, the prosecution can see patterns the defense cannot.
France took international heat for this. We never had to. Almost nobody knows the letter exists.
The reconstruction that proves the point
The wall is policy, not physics.
In 2020 a research team published JUSTFAIR: nearly 600,000 federal sentencing records from 2001 through 2018, linked to the judge who imposed the sentence, assembled from Commission files, Federal Judicial Center data, PACER docket initials, and the biographical directory of Article III judges. JUSTFAIR 2.0 extends the link through fiscal year 2023 — more than 140,000 additional records, 969 judges.
They did not hack the Commission. They joined public sources the judiciary already publishes in pieces that do not touch.
That is the tell. If a handful of academics can reconstruct judge-level sentencing from documents the government already releases, the 1988 letter is not protecting a secret that cannot be known. It is preventing the official file from saying what the unofficial file already can. An institution that forces researchers to glue the scoreboard back together out of docket numbers is not confused about measurement. It is declining to put its name on the measurement.
RAND and others, working from the reconstructed data, have since shown what the official file would have shown decades ago: racial disparities in federal sentences are large, and they vary considerably across individual judges. That finding is not a vibe. It is what happens when you finally attach a name to a sentence.
What the 2017 wording got wrong
The design I should not have written down the way I wrote it is this: I had demographic data — race, class, religion, location — circulating alongside the defendant's score at the time of booking.
Nothing about a defendant's race belongs near the calculation of that defendant's sentence. A system that computes a recommendation in proximity to those fields is one configuration change from doing something monstrous.
What I was reaching for, and stated badly, is that the data has to be retained. Attached to the case. Sealed through trial and sentencing. Aggregated afterward against the judge.
That distinction is the whole architecture. The judge sees the offense, the record, and the baseline. The judge never sees the demographics as inputs to the sentence. Months later the aggregate does: how many black defendants, how many poor defendants, how many defendants from this zip code came before this judge, charged with what, and what each received compared with everyone else charged with the same thing.
Blind at the moment of decision. Visible in the pattern.
I had the second half. I wrote the first half in a sentence a hostile reader could take the wrong way. The architecture is the correction. The 2017 thesis is not.
Data is disinterested. Choosing what to count is not.
I also wrote that quantitative data is, by its nature, disinterested. A batting average has no opinion about whether the batter succeeds. That sentence is still true.
Batting average was also the wrong statistic. That was Billy Beane's actual insight. Not that the number was false or badly calculated — batting average did what it claimed. The mistake was that baseball had spent a century counting the wrong thing. An inherited habit had dressed itself up as a simple observation. The data was disinterested the whole time. The choice of what to count was a value judgment nobody recognized as one.
That distinction has since been settled in a way that should end a certain argument permanently.
In 2016 ProPublica reported that a widely used recidivism tool wrongly flagged black defendants as future criminals more often than white defendants. The vendor said the tool was calibrated: a given risk score meant the same actual likelihood of reoffending regardless of race.
Both were correct. This was never a case of one side having better data.
Then Kleinberg, Mullainathan and Raghavan, and independently Chouldechova, proved the thing that removes the dispute from the realm of taste: when two groups have different underlying base rates, no classification system can satisfy both conditions at once. Not with a better algorithm. Not with more data. Not with a smarter team. It is impossible. Someone has to choose which kind of error to inflict on which group.
Every deployed system of this kind therefore contains a value judgment. Not a bias hiding in the file — a choice, made by somebody, about whose harm counts for less.
The only question is whether that choice gets made in public by people who can be named and argued with, or quietly, inside a vendor's model, by people who will never have to defend it.
Which is the argument about judges, one level up.
What I am not proposing
I want no fog here, because this is where the field went wrong and I do not want to be counted among them.
I am not proposing predictive risk assessment. Not for sentencing, not for bail, not for parole. Tools that forecast who will commit a future crime deserve the skepticism they now get, and the impossibility results are exactly why. There is no fair version — only a version where someone has decided in advance whose false positives are acceptable.
The distinction is not a technicality.
There is no fact about whether a given sentence was correct. There is a fact about whether a judge gave longer sentences to black defendants than to white defendants charged with identical offenses. The first is a prediction and cannot be made fair. The second is arithmetic.
I drew that line in 2017. The mathematics landed on it.
Prediction is a dead end for this problem and should stay dead. Measurement is untouched. Every serious objection raised against algorithmic sentencing in the last decade applies to forecasting behavior. None of it applies to counting what already happened.
The plea bargain
I argued in 2017 that plea bargaining is a subjective vehicle of injustice built into the system — unregulated, unevenly offered, serving convenience rather than justice. I have not changed my mind.
In a decent system, if a person committed the offense, the sentence is the sentence. If the state cannot prove it, there is no conviction. Everything in between is a discount for cooperation, connections, or the ability to hire someone who knows which prosecutor takes which deal. Administered without standards, without oversight, and without a public record of who was offered what.
In 2023 the American Bar Association's Plea Bargain Task Force — prosecutors, judges, defense attorneys, academics — found substantial evidence that innocent people are coerced into pleading guilty, that the practice worsens racial inequality through charge-stacking, and that defendants who exercise the right to trial face sentences commonly seven to nine years longer for having done so.
Ninety-eight percent of federal convictions come from pleas. In the states the trial has become a rumor. The Supreme Court has said outright that ours is a system of pleas, not of trials.
We do not have a criminal justice system. We have a negotiation system with a courthouse attached.
The strongest objection is capacity. Remove the discount and almost nobody pleads, and a system that tries two percent of its cases cannot try eighty percent of them.
That is true, and it is the confession. A system that would collapse if defendants used what the Constitution promised them is not a justice system. It is a machine that runs on people being too frightened to use the right.
I know abolition will not happen on a timetable I will live to see. That is not a reason to pretend the instrument is clean. It is the place subjectivity and expedience were installed on purpose. Measuring judges while leaving the discount unmeasured is finishing half the job again.
What is different now
The Herald-Tribune investigation took a team — reporters, data specialists, developers — and the better part of a year. They had to build the database themselves. The analysis was trivial. Turning documents into data was the entire project.
That barrier is gone, and it matters differently at each level.
Federally, the model is not even required. The work is done. The Commission receives every sentencing packet within thirty days and codes it. The database exists — complete, current, structured. Producing judge-level disparity analysis for the federal system is not a research project. It is a query. JUSTFAIR already proved the join is possible from the public side. The official file could do it in an afternoon.
What stands between the public and that query is not technology, cost, or difficulty. It is a letter from 1988 that nobody has had to read aloud in a long time.
At the state level, where most people are actually prosecuted, the records are public and unstructured. That was the wall. Extracting structured information from unstructured documents at scale is the thing current language models do well. What took a newsroom a year can now be done for an entire state by one competent person with public records and an API key.
When measurement was expensive, its absence could be blamed on budget. It cannot be now.
Every remaining barrier is a choice somebody is making.
What it costs to look
I said at the outset that the half which helps decision-makers gets built and the half that measures them does not. Courts are the clearest case. They are not a special case.
Hospitals, universities, agencies, employers, insurers — each is now choosing between a version of this technology that makes their decision-makers better informed and a version that shows the public how those decision-makers perform. The first will arrive quickly. The second requires the people being measured to approve their own measurement.
The pattern does not need a conspiracy. Nobody took a bribe. Nobody met in a garage. This is what capture looks like when it is procedural — people with the authority to decide whether they can be examined decide that they cannot, in the language of privacy and administrative discretion, and it never has to be defended because it never becomes a question.
A conspiracy can be exposed. This just sits there.
For courts, the case has not changed since 2017. We are not going to eliminate every unjust sentence. That was never the standard. The standard is whether a system that can measure itself is better than one that cannot, and whether a democracy is entitled to see how its judges behave.
We built cameras to settle arguments about a game.
We still pretend we cannot count what happens to a person.
The original 2017 paper remains on articles.swhammond.com, unedited, including the sentence I would now write differently.
Sources
The proposal and its origins
- Josh Salman, Emily Le Coz & Elizabeth Johnson, "Florida's Broken Sentencing System" — Herald-Tribune
- 28 U.S.C. § 994(w) — reporting requirements for sentencing courts
The tool that got built
- U.S. Sentencing Commission, Judiciary Sentencing INformation (JSIN)
- U.S. Courts, "Judiciary Studies Use of Online Tool in Presentence Reports" — the Federal Judicial Center randomized pilot
- Carlton Fields, "Sentencing Commission Data Tool Is Deeply Flawed" — the case against JSIN
France
- ABA Journal, "France bans and creates criminal penalty for judicial analytics"
- Slate, "France Has Banned Judicial Analytics to Analyze the Courts"
- Legal Futures, "French law banning naming of judges will not stop analytics"
The American version
- U.S. Sentencing Commission, Interactive Data Analyzer — see the data disclosure policy
- U.S. Sentencing Commission, Public Access to Sentencing Data and Documents (PDF) — contains the 1988 confidentiality agreement
- Sentencing Resource Counsel, SRC Dashboards FAQ
Fairness and prediction
- Alexandra Chouldechova, "Fair Prediction with Disparate Impact: A Study of Bias in Recidivism Prediction Instruments" (2017)
- Jon Kleinberg, Sendhil Mullainathan & Manish Raghavan, "Inherent Trade-Offs in the Fair Determination of Risk Scores" (2016)
- Richard Berk et al., "Fairness in Criminal Justice Risk Assessments: The State of the Art"
Plea bargaining
- American Bar Association, Plea Bargain Task Force — Fourteen Principles and a Path Forward (2023)
- Vera Institute of Justice, "In the Shadows: Plea Bargaining"
- Pew Research Center, federal trial and acquittal rates
The Images



























