Flucta

Closing the efficiency gap between algorithms and hardware

Top-DownProbabilistic Computing.

Modern AI tolerates extreme sparsity and stochastic sampling without losing intelligence. Conventional accelerators, built to stream dense tensors, turn that flexibility into unnecessary memory reads and power. Flucta builds the full stack that closes the gap, from algorithms down to our own silicon.

250:1

Energy to read a byte vs. compute on it

0

Retraining or fine-tuning needed to deploy

No loss

of intelligence, end to end

100×

Target intelligence per watt with the SPU

Reading a byte costs 250× more than computing on it.

On a modern accelerator, reading one byte from HBM costs roughly 25 picojoules; computing on it costs about 0.1. Models are increasingly tolerant to sparsity: in attention, in experts, in long contexts, a small fraction of the bytes carries the answer. Dense hardware cannot act on that tolerance, so every token pays for the full read.

Memory

HBM · hundreds of GB

KV cache, weights

bytes the model needs

Memory interface

Bottleneck25 pJ / byte

Compute die

0.1 pJ / byte

Waits on memory most of every token

Bytes the model needs

a small fraction

Bytes dense hardware reads

all, every token

250×

move compute

Widening the bus and stacking more memory does not remove the overhead: dense hardware still reads bytes the model never needed.

Flucta builds the stack that exploits the tolerance: probabilistic computing, sparsity, asynchrony, and approximate computing, run as one design from the algorithm down to the hardware.

Same intelligence.A fraction of the reads.

Read
Reads / tokenthe whole context

Legacy · Dense Read

Reads everything, every token

VS

Read
Reads / tokena small fraction

Flucta · Sampled Read

Reads only what carries signal

From algorithms to FPGAs to purpose-built silicon.

The speedups run as software on today's GPUs, with no retraining. The same kernels are moving onto FPGA prototypes now, and measurements from that hardware drive the Stochastic Processing Unit: purpose-built ASICs first, then 3D-stacked memory-on-logic. Every stage runs the same probabilistic abstraction on hardware with less overhead than the stage before.

Where we’re going

Today

End-to-End Probabilistic Algorithms

Running on GPUs today · FPGA validated · ICML 2026

2027

2D multi-chip SPU

28 nm · custom memory hierarchy

2028

3D-stacked SPU

Near-memory compute · datacenter

Physicists and chip architects, building from first principles.

Flucta's founders started the field of probabilistic computing with p-bits, and many of the field's Nature-level firsts are theirs: the first demonstration of 1,000,000 p-bits, the first carbon nanotube computer, and firsts in 3D chips and in hardware/software co-design.

Kerem Y. Çamsarı

Co-founder & CEO · Professor, UCSB (ex-Purdue)

  • Co-inventor of the p-bit (Nature 2019)
  • Landmark Nature paper with 2025 Nobel Laureate John Martinis
  • $25M+ federal research as PI
TS

Tathagata Srimani

Co-founder · Assistant Professor, CMU (ex-MIT/Stanford)

  • First carbon nanotube computer (Nature 2019)
  • Co-PI, $40M DoD AI hardware program, 40+ chip tapeouts
  • Led lab-to-fab transition of carbon nanotubes

Founding team

* Consulting
KP
Silicon

Kartik Prabhu

Ph.D., Stanford

SC
Algorithm / Kernel

Shuvro Chowdhury

Ph.D., Purdue

CD
Algorithm / Kernel

Corentin Delacour*

Ph.D., Univ. of Montpellier

KS
Algorithm / Kernel

Kemal Selçuk

Ph.D., UC Santa Barbara

NS
Algorithm / Kernel

Nihal Sanjay Singh

Ph.D., UC Santa Barbara

KB
Silicon

Kai Bartolone

M.S., Stanford

KC
Silicon

Kevin Callahan-Coray

Ph.D. Candidate, UC Santa Barbara

ZH
Inference Serving

Zengxiao He

M.S., Stanford

Advisors

SM

Subhasish Mitra

Stanford

Professor of EE & Computer Science · Robust Systems Group

SL

Suk Hwan Lim

ex-Google

Former Corporate EVP, Samsung Semiconductor

Funded by investors
who back foundational technology.

South Park Commons
Moxxie Ventures
Humba Ventures
Pear VC
Defined

Frequently Asked Questions

Flucta builds probabilistic computing across the full stack. Our algorithms and runtime speed up inference end to end on today's GPUs while keeping the model's intelligence intact, and the same principles carry into FPGA prototypes and purpose-built silicon.

Most hardware companies treat algorithms as static benchmarks, while software teams treat silicon as an unchangeable black box. Flucta takes a top-down, full-stack approach. By co-designing software algorithms, runtime SDKs, and physical execution targets around the principles of probabilistic computing, we capture orders-of-magnitude efficiency gains across the entire stack; benefits that isolated software or pure-play chip efforts cannot reach on their own.

Not initially. Though our principles will naturally lead us there in terms of custom ASICs, 3D-CMOS, and near-memory computation, our initial focus operates at the system and architectural level. Instead of waiting on novel physical devices, we re-architect memory hierarchies, routing logic, and execution pipelines on standard silicon processes to natively support sampling and probabilistic operations today.

Enormous efficiency is lost between the design abstractions of the modern stack. Algorithms are written as if moving memory were free, runtimes assume dense and deterministic execution, and silicon is built for worst-case precision. Each boundary hides what the layer below actually needs, and the waste compounds. Flucta removes these boundaries with a top-down design philosophy, carrying one probabilistic abstraction from the algorithm through the runtime into the silicon, so the gains at each layer multiply.

Email founders@flucta.ai to request the technical brief. We are hiring first-principles thinkers across physics, computer science, computer engineering, kernel development, and inference optimization.