a letter, for friends who think about this professionally
The disempowerment hedge
Dear friend,
I've started Caletta because I'm worried about what happens when people lose their economic bargaining power to AI. I'd like your help testing whether the thing I'm building could make a difference.
A king needed soldiers; a factory needed workers. States still need clerks and taxpayers. Those dependencies have given people some bargaining power, even when nobody intended to share it. AI could weaken that dependence.
I'm concerned about a world where alignment mostly works, but capability belongs to so few people that everyone else's influence steadily declines. Kulveit et al. call this gradual disempowerment. Drago and Laine's intelligence curse approaches the concern through economics: a state that depends less on its citizens' labor may have less reason to invest in them. Neither story needs a single dramatic turning point.
Caletta is a bet that people can retain more influence by owning useful AI capability themselves. I want to know whether the mechanism can support that ambition.
Where the bet could fail
There are several worlds where it would do very little.
If takeoff is fast and decisive, with a lasting gap between the frontier and everything else, distilling last month's capability onto local hardware may be economically irrelevant. The infrastructure could work and still fail at the reason I built it.
Private inference and training could still be valuable in that world. People have reasons to keep their work private even when their local model is weaker. But privacy alone would not restore the bargaining power lost when their labor became unnecessary.
Faster concentration could increase demand for private AI while making the broader counterweight less effective. That gives the business a possible source of demand even if the larger thesis fails. I don't want to confuse the two outcomes.
Caletta also depends on other people's work on misaligned takeover. It is not an answer to that problem.
The bet needs a meaningful diffusion window: frontier capability remains available to rent, small models can learn useful narrow tasks from it, and local hardware is affordable enough to run them. I think that window is open now. I don't know how long it will last, or whether people and small organizations can accumulate enough capability before it closes.
How the learning would work
The claim I most want tested is this: using a frontier model can produce the training material for a smaller model you own.
A record of real API work contains examples of tasks and their outcomes. Software on the user's machines can retain that record with permission, fine-tune a small model, and route tasks locally once it has shown it can handle them. I call the share served locally the Sovereignty Ratio. The test is to unplug and measure that share on the user's actual work.
The larger hypothesis is that ordinary use could distribute some of the capability concentrated in the labs. If the interaction record is useful training material, each customer has a way to retain part of what they paid to use. Pricing, rate limits, and provider terms could restrict that channel. How much useful distillation remains possible under those constraints is central to the bet.
This mechanism depends on the distillation channel staying useful. It does not require augmented humans to remain competitive with frontier systems at every task.
Possible breaks are easy to name: the record may be a poor curriculum, the small model may lack the capacity, or the economics may arrive too late. I'd like evidence about which constraint is most likely to bind.
The runtime and proxy run on Apple silicon. Two machines fine-tuned independently, exchanged adapters over a real network, and merged to byte-identical results for three rounds using a merge rule designed to tolerate corrupted contributions. The result was published with its failures. That was two machines, not a group; coordinated multi-node training is in build. The contribution verifier has one measured result against one defined corruption. The claim ledger separates measured and pending work, and the north star gives the full technical account.
What changes when capability moves offline
The proliferation objection deserves a direct answer. The intended scope is a small model fitted to one person's existing work, using outputs they could already obtain from a frontier teacher. That differs from releasing the frontier model's weights. The argument depends on that scope and on the smaller model's capacity limits; those assumptions need testing.
The model moves onto hardware that can run offline, without metering, observation, or revocation by the provider. That transfer of control is the product's value, and it also changes the risk.
Distillation also doesn't carry safety properties along with it. Fine-tuning even on benign data is known to degrade refusal behavior, so a distilled task model can't be assumed to keep its teacher's guardrails, and nothing in this design is allowed to lean on that assumption.
The training record need not consist of benign work. A motivated user could deliberately collect borderline outputs and train on them locally, removing the monitoring and rate limits that applied to the original requests. This is the refusal-laundering objection.
My argument is that the training material comes from what the teacher can be induced to emit. Even within that scope, removing observation and revocation matters. I cannot treat the provider's filtering as a safety mechanism under my control.
My commitments are to the user's control, an inspectable system, and the ability to leave. If a concrete case showed task-scoped distillation producing uplift beyond the teacher's accessible capability, the design would have to change. I also won't build a system that requires my approval to keep using a model you already hold.
Could the learning accumulate across people?
The Intelligence Commons takes the idea further: operators exchange and merge adapters, checking contributions before accepting them. It asks whether learning across many private loops could become a shared counterweight to concentrated compute.
That is speculation. A useful result on two machines does not establish that learning can accumulate across strangers, different tasks, and successive frontier releases.
The failure modes are listed on that page in the order they scare me. Poisoning is first. An open federation accepts training contributions from strangers, and unless a contribution can be checked cheaply before it merges, one motivated attacker corrupts what everybody shares. The verifier work above is aimed at exactly that. One measured result against one defined corruption is nowhere near an adversarially hardened gate.
What I'm asking of you
- Give me the strongest version of "this doesn't matter." The hardest objection for me is that the frontier's lead may grow faster than local models can learn, leaving the local share too small to matter. If you have a sharper form of that, or a better objection entirely, I want it in writing.
- Break the mechanism at a specific link. Curriculum quality. Small-model capacity. Merge stability under heterogeneity. Verifier hardness. Unit economics. Pick one and show me the measurements or assumptions that make it fail.
- If the thesis survives you, consider working on it. The work is to test this counterweight, with the exit test as its acceptance criterion and the same commitments binding me. The disagreement is worth having even if you never touch the code.
Banks's Culture is an appealing outcome, but the kindness of its Minds is part of the premise. I don't know how to rely on that premise in the world we're building. I'd like people to have some control of their own.
I'm asking for criticism and, if the work interests you, collaborators. There is no fundraising ask in this letter.
Warmly,
Travis