We're trying to solve the hard problems.

We think intelligence does not have to be one enormous model in a distant datacentre. It can be a society of small specialists, each expert at one thing, orchestrated and composed on the device in front of you. Our research runs from the single specialist, to the colony that coordinates them, to the compression that makes each one small, to the physics that could run them for a fraction of the power.

Fern, the specialist

Fern is the unit of our approach: a small model trained for one job. It runs on the device in front of you, a phone, a watch, an embedded chip, fully offline, and it answers in well under a second. It does not write a paragraph about what to do. It returns the action itself, the exact function call, intent turned into a command your software can run.

New Now shipping as Fernfly: Fern's research, productized for developers.
  • One job, done well: a specialist, not a generalist stretched thin
  • Intent to action: returns the exact function call, not a description of it
  • Sub-second responses, fully offline, on a phone, watch, or embedded chip
  • Private by construction: on-device means no inference logs and no data leaving the hardware
  • Suited to healthcare, defence, and legal work, where data cannot leave the building
  • Substantially fewer parameters, with quality measured and reported

At a glance

On-device
Runs offline on a phone, watch, or embedded chip
Sub-second
Fast enough to sit inside the loop
One job
A specialist trained for a single task
Function calls
Returns the action, not a paragraph about it
No logs
Data never leaves the device
Private
By construction, not by policy

The Colony

No single ant is smart. The colony is.

Our bet is that intelligence can come from a society of small specialists rather than one large generalist. Many Ferns, each expert at one narrow thing and limited on its own, talk to each other, route and hand off work, and compose into something that behaves like one capable mind, at a fraction of the energy and cost, and all of it running locally.

The name is not a coincidence. Antelligent is ant plus intelligent, and our logo is a colony. An ant colony is the canonical example of emergent intelligence from many simple agents: no single ant is smart, the colony is. The Colony is both what we call this research and the metaphor that drives it.

This is the research that sits underneath the Everyday Series, our product for orchestration and governance. The hard questions get worked out here. The product ships there.

Open research questions

Learned routing and dispatch

Which specialist should handle a task, and when to escalate to a larger model.

Composition and hand-off

Passing state between specialists without a giant shared context, the thing that makes big-model orchestration so token-expensive.

Where small beats large

Whether a society of small models can match or beat one large model at equal cost, and where that crossover falls.

All of it on-device

Doing this coordination at the edge rather than in a datacentre.

Active research directions

Higher ratios on reasoning architectures

Pushing compression further on next-generation model families without quality regression.

Provable quality bounds for safety-critical use

Formal quality guarantees for regulated domains such as medical and defence.

Vision-language and multimodal coverage

Extending our compression pipeline to protein, diffusion, and vision-language models.

Compression-native training pipelines

Building compressibility into the model from the start rather than applying it post-hoc.

Compression Refinements

Our compression technology already produces models with substantially fewer parameters, with quality measured and reported. But we're not done. Our research keeps advancing: broader model coverage, and compression techniques that are provably safe for regulated environments.

Active directions include architecture-aware pruning for next-generation model families, provable quality bounds for safety-critical applications, and compression-native training pipelines that build compressibility into the model from the start.

These improvements feed directly into our products: better compression means smaller on-device footprints, lower enterprise serving costs, and higher-quality compressed models in our open-source releases.

The Substrate

The furthest-out of our bets is about the physics of computing itself. Instead of a GPU calculating an answer step by step, let a set of coupled physical oscillators settle into one. You pose the problem as a physical system, let it relax toward its lowest-energy state, and read the answer off where it lands.

The interesting behaviour sits at the edge of chaos, the narrow band between rigid order and noise where a physical system is most expressive. Pair that with compute-in-memory, where storage and arithmetic are the same physical device rather than numbers shuttled back and forth, and the aim is orders of magnitude less power than a conventional chip.

We are honest about where this stands. This is early research, a bet on a direction rather than a product, and one path among several we are watching. We call early research early, and we publish the numbers that go against us alongside the ones that do not.

What we are exploring

Settling, not calculating

Coupled oscillators relax into an answer instead of a processor computing one step by step.

The edge of chaos

The expressive band between rigid order and noise, where a physical system carries the most information.

Compute-in-memory

Memory and arithmetic in the same device, so a value is not carried across the chip to be worked on.

Orders of magnitude less power

The reason to try at all: the target is a physical process that costs a small fraction of what silicon does today.

Interested in working with us on these problems?

We're looking for hardware partners, academic collaborators, and engineers who want to work on the foundations of sustainable AI, from the single specialist to the colony to the silicon underneath.