RL post-training infrastructure — the numerics and evaluation integrity of GRPO-style training. I find the numbers that stopped meaning what they say, in vLLM, TRL, verl and my own harnesses, and fix the ones I can reach. Background: inverse problems and medical-imaging AI at the University of Chinese Academy of Sciences, which is where the habit of measuring the ceiling before building the method comes from.
- vllm-project/vllm#55634: an aborted request's prefill leaks into
prompt_tokens_total; two contributors opened fixes the next day; one of them is still in review. - Four of the pull requests are merged, two of them code fixes in MLEvolve (#1 on MLE-bench). Six packages on PyPI — the four measured below plus groundwork-research and doubleblind-audit, both published 2026-09-24 — about 940 installs a month with mirrors excluded across the four (pypistats, read 2026-09-24 — the figure including mirrors is more than three times larger, and quoting that one would be quoting bandersnatch).
- batch-logprob-gap: bf16 batch-shape noise measured on eight models, fp16 shown to remove it from 1.5B up (not on the smallest), and a six-arm, two-seed GRPO run in which no arm moved reward by more than 0.013, half a standard error.
- huggingface/trl#7269 (closed unmerged, reopen requested):
removes the truncated-support term from GRPO's vLLM importance-sampling ratio — 0.896 → 1.001 at
top_p=0.8for an unchanged policy, and the logged mismatch 0.043 → 0.007 in a real colocated run.
Open to research-engineering work on RL post-training infrastructure — chengguo24@mails.ucas.ac.cn.
Each found by reading the source, each submitted with a reproduction and a before/after table. Four of the pull requests are merged. The recurring defect: a system keeps running while a number inside it quietly stops meaning what it says — nothing raises, nothing is red, and the number is wrong.
| Where | What | Status |
|---|---|---|
| vllm-project/vllm — 92k stars | #55634 an aborted request's prefill (+3,010 tokens over five aborts) lands in prompt_tokens_total while every per-request histogram stays at zero, so two billing gateways disagree by exactly the abandoned prefill · #56361 at n>1 eleven per-request metrics count the children and two count the request: a 311-token prompt at n=8 adds 2,488 to prompt_tokens_total for 360 tokens of prefill · #46530 a flaky(reruns=3) on the batch-invariance gate cuts a 41% chance of reporting a violation to 2.8% · #34333 a GSM8K regression gate 8.9 standard errors below its own baseline passes a five-point drop 99.5% of the time (script) |
fixes by two other contributors (#55837 closed, #55940 open, in review); #56923 mine, open |
| huggingface/trl — 19k stars | #6789 the defaults are safe for the truncation term but not per token: 7.5% of tokens leave the clip band at top_p=1.0, and the trainer's own chunking moves 8.0% of them with no inference engine involved at all · #7269 normalises the trainer's old log-probs over vLLM's replayed sampling support (return_sampling_mask), ratio 0.896 → 1.001 |
qualification posted; pull request closed unmerged, with tests and an end-to-end vLLM run — the measurement behind it is in batch-logprob-gap |
| verl-project/verl — 23k stars | #6280 the four-month question — is the on-policy log-prob mismatch a batch-invariance problem — measured on eight models: 1.8% to 60.0% of ratios leave [0.9, 1.1] in bf16, 0.00% in fp32, and only the MoE moves when the batch is merely regrouped (12.60%); PPO's clipped fraction holds while which tokens are clipped flips 1.7% to 10.1% |
measurement posted; flash-attn varlen is the one path it cannot test |
| NVIDIA/TensorRT-LLM — 14k stars | #19501 the RL-determinism RFC, three of its items sized from the trainer side: the batch=1 vs large-batch test it proposes finds 1.8% to 60.0% of ratios out of band in bf16 and none in fp32, on three cards; an fp32 LM head alone still leaves 1.5% to 32.2%, which is the RFC's own caveat, measured; replaying vLLM's sampling mask takes the truncated-support ratio from 0.896 to 1.001 | measurement posted |
| odlgroup/odl — 433 stars | #1730 ASTRA's 2d parallel geometry silently dropped a shifted detector (#359, open since 2016); found by a guard that requires every mismatch axis to change the sinogram — it changed it by 0.0 | open, review addressed |
| InternScience/MLEvolve — 445 stars · #1 on MLE-bench | #8 a default-on leakage guard compared floats with ==, so its 117 lines had never run · #9 the memory block sorted minimise-metrics backwards |
both merged, third contributor |
| github/awesome-copilot — 39k stars | #2854 research-harness-engineer, an agent definition for research as a falsification loop · #2938 a contributor-table row with one cell too many |
both merged, on the contributor wall |
| InternScience/InternAgent — 1.4k stars | #27 a task maximised test-set MSE, so a worse error was rewarded — one character | open |
| ResearAI/DeepScientist — 3.3k stars | #110 pytest aborted collection on a clean checkout, 56 test files lost to one undeclared import |
open |
| cline/prompts — 1.2k stars | #46 breakthrough-research-discipline, a rule file for research as a falsification loop: the ceiling and the tuned baseline before the method, and a claim whose polarity is stated before the result |
open |
RL post-training and evaluation integrity
| Project | One sentence and the number it stands on |
|---|---|
| batch-logprob-gap | In bf16 a trainer gives the same token a different log probability depending on its batch shape: 1.8% to 60.0% of importance ratios leave [0.9, 1.1] across eight models, 0.00% in fp32; fp16 scoring removes it on the Qwen2.5 models from 1.5B up but leaves 22% on pythia-160m; six GRPO arms on two seeds each land within 0.013 of the default on the same seed. Every table is regenerated from the JSON, and the six-page report is assembled from those tables by a script that refuses to build if a number in its prose is not in a committed result file. |
| breakthrough-harness | Make a research agent hard to fool: adapters for nine stacks, every README claim asserted by a test; works with DeepSeek Harness as is. |
| doubleblind |
An agent cannot check its own output: the reasoning that produced a claim is the reasoning being asked to check it. Three layers that are blind to different defects — re-derive every number, hand the artifact to a reader told nothing, audit the figure a reader will actually see. On a document that had already passed 37 mechanical checks and two rounds of its author's own review, a zero-context reviewer on a different model returned nine findings; 4 of the 5 that survived verification were correct numbers in sentences that did not follow from them, which no quantity-comparing check can see. A ledger of 23 real defects records which layer missed each one — and each layer ships a test asserting the defect it cannot catch. |
| groundwork |
The complete research pipeline for coding agents — direction, gate, experiment, claim, paper, submission — with the stage every other toolkit is missing: one that returns NO-GO. Four measurements before the first experiment refuse a direction whose ceiling, tuned baseline, random arm or positive control already answers it; a pre-registration version control can date; cluster tooling that declines a card whose free memory is somebody else's leftovers and refuses to split one arm across two GPU models. 43 stage notes, each written from a failure, and an archive of 10 ways a direction dies with the cheap test that catches each. The command-line side runs a plan overnight and stops at the first step that exits non-zero — the morning report says which steps therefore never ran — and sweeps every gate over a project, where n/a is printed as loudly as FAIL because a check that passed for want of anything to check has said the opposite of the truth. |
| ifeval-reproduction | A published IFEval score on one shared GPU, three arms, a pre-registration chain CI re-hashes on every push; the "11-point gain" of the third arm died to a paired test — the subsample was easier. |
| taichu-eval-reproduction | Two ZDTaichu5.0-9B model-card numbers re-measured on the card's protocol (thinking on, LLM answer extraction), on the full sets rather than a subsample: CV-Bench 87.26% [85.93, 88.51] against the card's 86.82, MathVista 82.40% [79.90, 84.71] against 84.50, both cards inside. The four estimators this page used to quote are superseded and reported as errors against the measurement — all high on CV-Bench, all low on MathVista, from the same subsample method. And the MathVista verdict rests on 16 truncated generations: 37 of 1,000 never closed their reasoning block, the extractor credits 16, and counting those as no-answer gives 80.80% [78.22, 83.20], which does not contain 84.50. |
Inverse problems, world models, tools
| Project | One sentence and the number it stands on |
|---|---|
| ct-reconstruction-harness | LoDoPaB-CT baselines from scratch: matched TV 33.00 ± 0.33 against a published 33.36, TGV at +0.70 ± 0.04 dB over it on 128 held-out images (125 wins); its first table quoted 16 images that run 0.9 dB easier than the other 112, and the README withdraws what rested on them; the seven undocumented layers the reproduction needed are each written down in DEBUGGING.md. |
| topocheck |
Five checks for topology-aware segmentation claims, including the random-repair baseline that beat every learned repair I tried. |
| worldmodel-from-scratch | A world model in an afternoon, then where it breaks; CI checks only the claims that hold on any machine. |
| world-model-map | A researcher's map of open-source world models — what each claims, what its authors say it cannot do, an evidence grade per entry; CI re-resolves every citation. |
| kakeya-conjecture-lab | An interactive Kakeya-conjecture lab whose dimension meter recomputes its own numbers in the test suite. |
| scholarcheck |
Verify citations before a reviewer does; figures that check their own layout; find what a converter silently dropped. On PyPI, installed by people I have never met. |
Stars are a poor signal at this size. The number that does not depend on trusting me is on PyPI — about 940 installs a month across four packages (2026-09-24; the fifth and sixth were published on the 24th and have no month behind them yet) — and clone traffic is below, with my own CI checkouts excluded: the first version of this table counted them as people, and the correction, with the raw API responses, is in MEASUREMENT.md.
Two things that figure is not. It is not the number PyPI's own badge shows: about 70% of the raw
downloads are mirrors (3,176 with them, 944 without, on 2026-09-24), and quoting the larger one would be
quoting bandersnatch rather than a person. And it is not the launch figure: the snapshot of
2026-09-08 in data/traffic.json sums to 2,222 and the one of 2026-09-24 to 944, which is
what happens when four packages published in mid-August roll their launch week out of a 30-day window. scripts/refresh_traffic.py rewrites both places the
total appears from the same snapshot that feeds the badges, fails if either sentence stops matching,
and dates the total by its oldest part when pypistats rate-limits a package.
| Repository | People who cloned it, CI excluded (2026-09-24) | Raw total | PyPI / month |
|---|---|---|---|
| batch-logprob-gap | — | ||
| scholarcheck | |||
| kakeya-conjecture-lab | — | ||
| topocheck | |||
| docxaudit | |||
| ifeval-reproduction | — | ||
| worldmodel-from-scratch | — | ||
| sciglyph | |||
| ct-reconstruction-harness | — | ||
| world-model-map | — | ||
| breakthrough-harness | — | ||
| taichu-eval-reproduction | — | ||
| doubleblind | — | ||
| groundwork | — |
Clone counts are as verified against the API on the date in the header; the automatic refresh needs a traffic token that this page does not carry, so the date advances only when they are re-checked by hand.
Most of the methods work is in unreleased repositories because a paper or a filing is still open: four lines on guarantees for medical image segmentation, what a pre-treatment image can establish about a treatment decision, image synthesis for adaptive radiotherapy, and prognostic markers in functional imaging — one patented, two with reviewers, one written and held. Each began with a measured ceiling and a random baseline before any method was built, and each keeps a written record of the attempts that did not survive those checks. Happy to go into any of it properly in a conversation.
Last verified against the API on the date above; this page could not re-check them today because no traffic token is configured.
Numbers go to disk before sentences are written about them. Every guard is run against a deliberately broken input and has to fail for the right reason before it counts. When a result does not survive that, the repository says so — the honest number is more useful than the flattering one.