Stop re-discovering dead links
Academic dataset hosts rot. A registry that records what actually worked — and when a path goes stale, gets corrected in place — beats re-searching from zero every time a benchmark needs to be re-run.
Research & Tooling
An open-source index of where the public tracking and perception benchmarks actually live — so re-downloading one later doesn't mean re-discovering a working source from scratch.
A pointer registry, not a dataset mirror.
Cross-benchmark validation is core to how Perception Origin measures its perception software — dwell/presence recall under crowd density, tracking robustness, and general multi-object-tracking research draw on public academic benchmarks (MOT17, MOT20, PETS2009, CrowdHuman, DanceTrack, the KITTI/nuScenes/Waymo family, and more). Several of the original publisher hosts for these benchmarks are dead, slow, or have moved — finding a working source took real effort each time. The Benchmark Dataset Registry is that effort, saved and kept current: fetch scripts, exact source URLs, and per-dataset license notes for 30 public datasets. The registry contains no dataset files itself — only the recipe to get each one, and an honest record of what's actually verified versus merely a reachable candidate.
Built for people and for the AI research tooling that increasingly does this work.
Academic dataset hosts rot. A registry that records what actually worked — and when a path goes stale, gets corrected in place — beats re-searching from zero every time a benchmark needs to be re-run.
Every entry is marked Verified, Candidate, Registration-gated, or Unconfirmed — no fabricated download paths, no glossing over datasets that require a license agreement the registry can't sign on anyone's behalf.
Increasingly, the thing consulting a registry like this isn't a person browsing GitHub — it's an AI agent (Claude, GPT, and others) planning a benchmark run. The registry is written so an agent can read a dataset's entry, use the recorded path, and — if that path is now dead — correct it in place rather than silently working around it in one session and leaving the next one to rediscover the same dead end.
Each dataset's real terms are documented — including datasets deliberately excluded for ethical reasons (see the registry's notes on why DukeMTMC isn't listed) and datasets that require registering directly with the publisher, for which no automated fetch script exists or will be added.
Public, no account required to read it.
halfwinddev/manasaview-benchmark-datasets — fetch scripts, source attribution,
and LICENSE-NOTES.md for MOT17/MOT20/PETS2009, DanceTrack, CrowdHuman,
SportsMOT, VisDrone, WildTrack, the KITTI/nuScenes/Waymo/BDD100K/Argoverse 2 family,
Ego4D/Ego-Exo4D, and more.