Open Weights. Open Artifacts. Open Licenses.
Open Data. Open Origins.
OpenWALDO.

AI is not open source without the source.

WALDO makes training data behave like open-source input: named, reviewable, versioned, attributable, and verifiable—and carries that provenance into the models built from it.

The working toolchain is being developed in public: core data and model workflows run end to end today, while contributors help refine the interfaces and expand what comes next.

THE AI

SOURCE
CODE.
OpenWALDO
01OPEN THE INPUTSTRACE THE OUTPUTS

The proven model

Open source communities
built the modern world.
Now let’s build AI together.

Open source gives people and organizations a common foundation they can inspect, improve, teach, and build upon. It turns users into contributors, competitors into collaborators, and shared problems into infrastructure that operates at massive scale.

01

A proven way to build at scale.

Linux, the web, cloud infrastructure, containers, supercomputing, Python, PyTorch, and much of today’s AI tooling grew through open communities working together.

02

Give AI a shared foundation.

One community corpus reduces duplicated foundational work and gives every team a stronger place to begin—while leaving each builder free to create, differentiate, and compete above it.

03

Let every contribution compound.

A source added once can support many models. A correction improves the public record. Better tooling helps every future contributor, and the benefits remain available to everyone.

The public commons

The community corpus.
Built to scale in public.

Individuals, researchers, and organizations build a shared training-data commons through one accountable public record. Git review governs meaning, content-addressed storage carries the bytes, and DCO sign-off keeps responsibility attached as the corpus grows.

Indexed training materialreference tokens across documents

Loading the live index breakdown…

corpora shards asserted license identifiersLoading live index…

Open means the sources

Weights are binary artifacts, not source.

Open source AI requires open training data and a verifiable path from inputs to every artifact.

01

Git governs meaning.

Small metadata stays readable, reviewable, and attributable.

02

Hashes govern identity.

Canonical Parquet objects are addressed by their content, not trust.

03

BOMs cross boundaries.

Resolved data and model lineage travel into runs and release packages.