About us · our thesis

Own intelligence. Kill the billable token.

Syzygy’s mission is to enable cost-efficient, outcomes-driven AI. We are a group of mathematics and physics researchers rebuilding AI deployment from the ground up as we think it should be: efficient, locally controlled, and owned rather than rented.

98%

of full-precision performance preserved by Mach-1 Small

160 tok/s

on devices with as little as 16 GB of unified memory

~1/10

the cost of the original deployment

<2 bits

per parameter: models and inference built for it

The problem

Renting intelligence is misaligned by design

Syzygy’s mission is to enable cost-efficient, outcomes-driven AI, and to kill the billable token. Today, inference providers of both open and closed-source models charge primarily for usage, whether per GPU hour or million tokens, rather than outcomes. Given the nondeterminism of token consumption per unit of work and the enormous energy, hardware, and labor costs of building and maintaining datacenters, this tradeoff is reasonable but painfully less than ideal. AI-native services mitigate this, aiming to more efficiently convert compute into outcomes by designing specialized products for specific industries. But that does not eliminate the upstream misalignment with compute providers, which is fundamental to the current model of “renting intelligence”.

We are already seeing the first signs in software engineering. Companies are spending extraordinary amounts on tokens approaching the cost of human labor itself, forcing them to cap token consumption per engineer. We predict this same tension will spread into any industry that begins to depend on AI.

Our thesis

Intelligence you own, like an operating system

Syzygy believes that cost-efficient, outcomes-driven AI becomes structurally feasible when users are able to own intelligence as easily as an Operating System. Once intelligence is owned rather than rented, we predict incentives will shift from metering units of reasoning to completing as much useful work as possible.

This is possible if the hardware and energy costs of AI begin to approach the cost of a modern personal computer. We see the bottleneck as four related problems:

capability / byte

Capability per byte of memory

capability / op

Capability per unit of computation

ops / watt

Computation per watt

work / dollar

Useful work per dollar of total system cost

These problems span the entire stack: model architecture, numerical representation, inference software, memory systems, and chip design. Syzygy is beginning with two foundational technologies:

models

Sub-2-bit language models

engines

Inference engines, and eventually chips, designed specifically for sub-2-bit computation

What we're building

Starting with sub-2-bit models and inference software

Most modern models and accelerators were built around relatively high-precision arithmetic. But inference does not require every parameter to be stored and processed at that precision. If model capability can be preserved below two bits, the cost of storing parameters, moving them through memory, and computing with them can fall dramatically.

We are starting with sub-2-bit models and inference software. Our proprietary compression algorithm and inference engine bring the capabilities of large language models to small, locally controlled devices. Our first model, Mach-1 Small, preserves 98% of the measured performance of its full-precision source model while supporting large contexts on devices with as little as 16 GB of unified memory. It runs at up to 160 tokens per second and approximately one-tenth the cost of the original deployment.

In the coming months, we plan to release Mach-1 XS, a smaller 1-bit model designed to run on mobile devices, and Mach-1 Medium, which brings the capabilities of a 120-billion-parameter model to laptops with 36 GB of unified memory.

Over the longer term, we plan to build affordable integrated systems capable of running large models entirely within enterprise environments.

Mach-1 Smallavailable now
Preserves 98% of the measured performance of its full-precision source model, with large contexts on devices with as little as 16 GB of unified memory.
Mach-1 XScoming months
A smaller 1-bit model designed to run on mobile devices.
Mach-1 Mediumcoming months
The capabilities of a 120-billion-parameter model on laptops with 36 GB of unified memory.
Integrated systemslonger term
Affordable integrated systems capable of running large models entirely within enterprise environments.

The future

The future of AI will be heterogeneous

The future of AI will be heterogeneous. Some intelligence will run in hyperscale data centers. Some will run inside companies. Some will run on laptops, phones, robots, vehicles, and machines that have not yet been built.

Not every model will run locally, and not every workload should. The important question is whether people and companies have a genuine choice between renting intelligence and owning it.

We could not be more excited to launch Syzygy.

We are a group of mathematics and physics researchers rebuilding AI deployment from the ground up as we think it should be: efficient, locally controlled, and owned rather than rented. If that mission resonates with you, we invite you to join us.