A proven way to build at scale.
Linux, the web, cloud infrastructure, containers, supercomputing, Python, PyTorch, and much of today’s AI tooling grew through open communities working together.
AI is not open source without the source.
WALDO makes training data behave like open-source input: named, reviewable, versioned, attributable, and verifiable—and carries that provenance into the models built from it.
The working toolchain is being developed in public: core data and model workflows run end to end today, while contributors help refine the interfaces and expand what comes next.
THE AI
SOURCEThe proven model
Open source gives people and organizations a common foundation they can inspect, improve, teach, and build upon. It turns users into contributors, competitors into collaborators, and shared problems into infrastructure that operates at massive scale.
Linux, the web, cloud infrastructure, containers, supercomputing, Python, PyTorch, and much of today’s AI tooling grew through open communities working together.
One community corpus reduces duplicated foundational work and gives every team a stronger place to begin—while leaving each builder free to create, differentiate, and compete above it.
A source added once can support many models. A correction improves the public record. Better tooling helps every future contributor, and the benefits remain available to everyone.
The public commons
Individuals, researchers, and organizations build a shared training-data commons through one accountable public record. Git review governs meaning, content-addressed storage carries the bytes, and DCO sign-off keeps responsibility attached as the corpus grows.
Loading the live index breakdown…
Open means the sources
Open source AI requires open training data and a verifiable path from inputs to every artifact.
Small metadata stays readable, reviewable, and attributable.
Canonical Parquet objects are addressed by their content, not trust.
Resolved data and model lineage travel into runs and release packages.
Enter the project
Start from a blank architecture or supported open weights. Forecast, train, inspect, and export without dropping the lineage.
See the model lifecycle ↗ 02 / CONTRIBUTINGTurn local material or reviewed acquisition recipes into auditable, Git-governed corpus contributions.
See how contribution works ↗