The long-term direction
The north star
CanonicalClaims elsewhere derive from this page. Updated 2026-08.This is Caletta's long-term direction. The odometer records what is running, measured, or still planned, along with the company's commitments. You can also read the vision in plain words or see the deployment plans.
The mission and the vision
What all of it is for
Mission: Turn the intelligence people rent into capability they own.
Vision: A world where anyone can run frontier-grade AI on hardware they hold, and prove it still works when they unplug from the labs.
The ambition is to keep humans free. Owned means: it runs within a custody domain you control, it learned from work you consented to use, it keeps working after you unplug from the frontier labs, and the share of your workload it covers goes up over time. A custody domain can be a laptop, workstation, on-prem cluster, or datacenter capacity you control. The runtime, proxy, mesh, and security tools are being built to support that ownership.
Free from what
What "free" means here
"Once men turned their thinking over to machines in the hope that this would set them free. But that only permitted other men with machines to enslave them."
— Frank Herbert, Dune, 1965
Herbert's remedy was to ban the machines. Ours is to own them. My concern is that AI could let concentrated power stop needing the people it has power over. For Caletta, freedom has three practical requirements:
- Free from dependency. You don't rent your memory, reasoning, and tools from four companies you can't afford to leave.
- Free from surveillance. The intelligence you use isn't also the thing watching you, because it runs on your side of the line. This sharpens as agents land on the desktop: an agent needs your screen, files, and credentials to work, so vendor-side agents turn "what you chose to send" into "everything you do on your own machine." The observation layer belongs in the operator's custody domain or nowhere.
- Free from exclusion. Capability ships wherever the operator controls the hardware: a laptop, a workstation, their own servers, a datacenter fleet. Today the laptop is an expensive one; the floor falls every year and the stack is built to follow it down.
Freedom here is capability you can retain, inspect, export, and run without any vendor, including us. That is a thing you can build, because physical custody moves the leverage to your side of the line. It does not make cutoff impossible.
The test that makes it real
The unplug test
The acceptance test is:
Unplug from the frontier labs, and reproduce a declared fraction of your real workload, at a declared quality bar, from your own model, corpus, evals, and policy, on a clean machine, with no Caletta service required.
The Sovereignty Ratio is the fraction of a user's workload served within their own custody domain, above an explicit quality bar, after learning from their own consented frontier interactions. The domain includes every machine the user has enrolled and controls. Progress means increasing that fraction on useful work.
Report a task, not a company-wide score. Publish the hardware, numerator, denominator, quality threshold, evaluator, and held-out period. Use independent acceptance criteria. Blinded human pairwise review and objective task checks decide acceptance; a true outcome label outranks a judge when one exists. An LLM judge can help during iteration, but cannot approve the result. Declare the workload before the run. Use real, recurring work and report local serving costs. A high ratio on an easy selection of tasks, or at an impractical price, would not establish a useful exit.
The first Ownership Loop is deliberately narrow: capture a consented, encrypted corpus for one repeated task; distill a small model and router within the operator's custody domain; score them against the frontier baseline; and route qualifying calls locally. The v0 task is code-review triage. The company owns a permission-clean corpus, and later patches and tests provide outcome signals, so this test needs neither customer PII nor a design partner. v1 is a paying customer: a narrow deployment for runtime control and audit, followed by the same loop on the customer's queue. That second step tests the economics.
The ladder
The steps toward ownership
Each step needs evidence before it supports the next:
- You own your inference. Capable models run within a custody domain you control. RunningRunning today on Apple silicon: the local proxy and the
mlx-goruntime, with more of the stack built and queued for release. Personal machines are the first substrate we can measure, not the ceiling the design is built for. The proxy's position enables a second capability, in build: outbound-traffic screening, where a local model flags PHI, credentials, and privileged material before it leaves the custody domain. Necessarily local, since a cloud checker would itself be the leak. A screen that helps you catch, not a compliance guarantee. This position compounds as agents move onto the desktop, by structure rather than by anyone's bad intent: agentic work requires broad local access (files, repos, sessions, screen), whatever is read crosses to the vendor side because that is where the model runs, and retention and training terms are the vendor's to set and change, so the only durable control is a boundary the operator owns, which is what the proxy is. Ambient agent traffic is also where screening stops being a nicety: no human reviews what an agent sends. And every call the local model serves is a slice of the machine's activity that never crosses at all: a value that lands before the full loop is proven. - You own the loop. Consented model interactions become training signal inside your custody domain; machines you enroll can contribute without pooling the raw interaction corpus. What crosses any given link is a per-link setting chosen by the operator on each side, which is what lets custody domains compose into arbitrary topologies (solo, pair, cluster, federation) without a global sharing policy. The settings are defined below. In buildIn build, off a first cross-machine result: two Apple Silicon machines independently fine-tuned, exchanged, and robust-merged real LoRA adapters over a real network to byte-identical files in three consecutive rounds, validation loss falling from 6.07 to 2.23. That is the two-machine exchange-and-merge path, measured and published with its failures. Coordinated multi-node training is the next run, and the public claim ledger tracks what is measured versus pending. The transport is public (
mlx-go-iroh); the integrated training datapath ships next. - Anyone can own it. The stack is open, exitable, and inspectable, so others can run it, fork it, and extend it without us. By designConstraint by design: the three falsifiable lines.
- Everyone does own it. Owned capability becomes the default a person reaches for, the way the PC took computing off the mainframe. DirectionThe north star: claimed only as the direction the first three rungs point.
The Exchange Dial specifies what may cross each connection. It draws on Eric Hughes's definition: "Privacy is the power to selectively reveal oneself to the world" (A Cypherpunk's Manifesto, 1993). There are five levels: nothing (local only, no teacher, capability frozen at what you have); metrics (tasks, quality, and the ratio, without content); weight updates (adapters and deltas, without raw work); derived signal (de-identified or synthesized examples); and raw work (what a rented API ordinarily receives). The measured two-machine result used weight updates. These levels need separate privacy evidence: adapters can leak training data, so exchanging them is not equivalent to sending nothing. Each connection has both a send policy and an acceptance policy: from whom, and verified how. The sparse-recompute contribution gate belongs on the acceptance side. Both policies must be inspectable, with the most private default that still performs the job.
The learning loop is the decisive step. Local inference, privacy, and inspectable code are useful, but they do not by themselves reduce dependence on frontier reasoning. The Sovereignty Ratio tests whether learned capability actually takes over part of the work.
Where this sits in the landscape
How the pieces fit together
Other projects already supply many of these pieces. Ollama, LM Studio, and llama.cpp support local inference. Prime Intellect, Nous, and Pluralis have trained across real networks at scales beyond our small first result. Several labs release downloadable weights. There is also overlap in the motivation: Thinking Machines writes that "a single locus of value alignment, however well run, becomes a locus of power to be captured." The diagnosis and the individual components are shared ground.
Caletta aims to connect those pieces into a measured learning loop with an exit test. Rented intelligence produces an owned model, tested on the operator's work after disconnecting from the labs. We have not found another measured loop with this particular acceptance criterion. If others build one, the operator should benefit: the corpus, router, evaluations, and adapted weights remain theirs to carry between implementations.
One narrow technical claim, because anyone working in the field will check it: trustless verification of inference on an untrusted mesh has published, shipped approaches, including Prime Intellect's TOPLOC. Integrity of a training contribution across many steps is a different problem from recomputing one forward pass, and we have not seen a published, adversarially-tested answer for it. That is the leg our mesh targets, and it has a first measured result: a verifier that re-executes a sparse sample of a committed update caught 82 percent of a defined corruption while sampling under half a percent of the bytes, at about two percent of verifier time. Next is the same verifier against a real recomputer, then the end-to-end gate: a defined invalid contribution rejected while the ungated path accepts the same input. That proves the implemented gate, which is the claim we make, rather than general Byzantine-robust training, which we don't. The mesh matters as the training arm of your loop, not as a compute network racing the ones others already run.
The circle
No gatekeeper, including us
The design principle is that no gatekeeper decides who gets capability, including Caletta. The software should let people run it without asking us. "Keep humans free" is a direction; the concrete work is to avoid dependencies that would let us withdraw what someone already owns.
The rule, and its receipt
The rule that decides what ships
Every shipped decision either widens the circle of who owns capability, or it doesn't ship.
That selection rule guides decisions. The three prohibitions define observable violations a reader could point to. A decision log records features and deals turned down because they would concentrate control, including control in Caletta. It is kept from day one and shared as entries accrue, so the costs of following the rule can be checked.
What this asks of the people reading it
Investors, builders, everyone else
- Investors: the near-term case is a real, pre-revenue infrastructure company selling sovereignty to the people the mission is for. This page tells you the size of the thing that case is aimed at. If "keep humans free" reads as a reason to pass, we probably want different rounds; if it reads as the reason the company won't quietly sell out to a gatekeeper, we agree.
- Builders and collaborators: this is the recruiting pitch and the constraint at once. You'd be building the counterweight, under commitments that bind the founder too.
- Everyone else: the alternative exists, a usable, ownable other path that keeps widening the circle instead of narrowing it.