Harshil Chudasama
All projects

Case study

adduce

ML research artifact auditor

A command-line tool that asks one question of every number in an ML paper: can you point to the code and data that produced it, and would they produce it again on another machine? It reports what it finds and never claims a project is reproducible.

Role
Creator and maintainer
Year
2026
Status
Live
Stack
PythonTyperStatic analysisPyPIGitHub Actions
adduce

What it is

adduce is a command-line tool for machine learning researchers. It reads a project's code, configs, data, and results, and reports what someone would need in order to reproduce the paper's numbers. It runs locally, uses no network by default, and never runs the project's own code.

pipx install adduce
adduce check .

Version 0.2.0 is on PyPI. It is MIT licensed and in beta.

The problem

A paper reports 81.4% accuracy. Six months later a reviewer, or the author, tries to find the run that produced it. The config has changed since. Only some of the random seeds were set. The dependencies were never pinned, and a model downloaded from a hub was never locked to a version. Nobody did anything wrong. These gaps build up quietly and only show when someone tries to rerun the work.

Reproducibility checklists ask authors to confirm all of this from memory. adduce checks the repository instead.

How it works

  • 78 checks in 17 categories. They cover commands and scripts, dependencies, data, random seeds, numerical precision, downloaded models, and whether the paper's numbers match the logged results.
  • Specific findings. Each finding is marked pass, partial, fail, not applicable, or unknown. It points to the file and line where it applies and suggests a fix. Partial matters because most real projects half-meet most checks.
  • Only the checks that apply. A project with no notebooks is not scored on notebooks.
  • Drafts for reviewers. It drafts NeurIPS and ACL checklist items, an ACM artifact appendix, archive metadata, and a claim-by-claim evidence trail. All of it is marked as a draft for the author to confirm.
  • Extensible. Labs can add their own rules through a plugin API, and a GitHub Action runs adduce on pull requests.

A real run

On nanoGPT at commit 3adf61e, adduce 0.2.0 scores the project 54/100. It estimates 23 to 83 minutes before a reviewer gets a first result, and rates that Risky. Five of the fifteen categories that applied:

AreaScoreWhat it found
Code & Execution8/12Commands documented, but no run script
Environment & Tooling1/10No dependency list, lockfile, or container
Data8/10No checksums on the data
Determinism & Model3/12Some random seeds set, not all
Remote Artifacts & Rot1/4Hub downloads not pinned to a version

Findings point to exact lines, such as the TF32 setting at train.py:107 and an unpinned model download at model.py:238.

When a project lists its claims, adduce traces each one from the sentence in the paper to the logged number behind it. If the paper says 81.4 and the log says 81.37, the claim is marked partial for a person to check.

Design choices

Offline and read-only by default. adduce does not touch the network or run the project's code unless asked. Running untrusted research code just to inspect it would be the bigger risk. The cost is that settings built at runtime can't always be resolved, and those come back as unknown.

Honest about its limits. A score next to the word "reproducibility" reads like a verdict, and static analysis can't support one. adduce reports what it detected and never says a project is reproducible. Its scores are labelled experimental until they have been checked against human reviewers.

Stable output. The same repository gives the same report on every run, so it can gate CI. Release 0.2.0 added a benchmark suite to CI. It fails the build if a repeated run stops matching byte for byte, a synthetic project's score moves, or the tool starts reading files more often.

Testing

adduce is tested against 15 real research repositories pinned to exact commits, and 14 small synthetic ones built to trigger known results. The synthetic ones run with the regular test suite. In 0.2.0, each file is read once per run instead of up to three times, which cut the largest test repository from 21.7 to 16.7 seconds without changing any test repository's report.

What I would do next

  • Check the scores against people. Until enough real projects have been reviewed by hand, a score is a rough ranking. That comparison has to happen before adduce can claim its scores mean what they say.
  • Read claims from papers more reliably. Pulling claims out of LaTeX is still in development. Work on the development branch reads more table formats and no longer mistakes a label such as recall@1 for a result.
  • Resolve more settings. Values assembled at runtime are where the analysis gives up and returns unknown.
  • Recover old downloads. Pinning locks a model version from now on. It can't tell what a reference pointed to two years ago.