- Co-inventor of the p-bit (Nature 2019)
- Landmark Nature paper with 2025 Nobel Laureate John Martinis
- $25M+ federal research as PI


Closing the efficiency gap between algorithms and hardware
Top-DownProbabilistic Computing.
Modern AI tolerates extreme sparsity and stochastic sampling without losing intelligence. Conventional accelerators, built to stream dense tensors, turn that flexibility into unnecessary memory reads and power. Flucta builds the full stack that closes the gap, from algorithms down to our own silicon.
250:1
Energy to read a byte vs. compute on it
0
Retraining or fine-tuning needed to deploy
No loss
of intelligence, end to end
100×
Target intelligence per watt with the SPU
Reading a byte costs
250× more than computing on it.
On a modern accelerator, reading one byte from HBM costs roughly 25 picojoules; computing on it costs about 0.1. Models are increasingly tolerant to sparsity: in attention, in experts, in long contexts, a small fraction of the bytes carries the answer. Dense hardware cannot act on that tolerance, so every token pays for the full read.
Memory
HBM · hundreds of GB
KV cache, weights
bytes the model needs
Memory interface
Bottleneck25 pJ / byte
Compute die
0.1 pJ / byte
Waits on memory most of every token
Bytes the model needs
a small fraction
Bytes dense hardware reads
all, every token
250×
move ≫ compute
Widening the bus and stacking more memory does not remove the overhead: dense hardware still reads bytes the model never needed.
Flucta builds the stack that exploits the tolerance: probabilistic computing, sparsity, asynchrony, and approximate computing, run as one design from the algorithm down to the hardware.
Same intelligence.A fraction of the reads.
Legacy · Dense Read
Reads everything, every token
VS
Flucta · Sampled Read
Reads only what carries signal
From algorithms to FPGAs to purpose-built silicon.
The speedups run as software on today's GPUs, with no retraining. The same kernels are moving onto FPGA prototypes now, and measurements from that hardware drive the Stochastic Processing Unit: purpose-built ASICs first, then 3D-stacked memory-on-logic. Every stage runs the same probabilistic abstraction on hardware with less overhead than the stage before.
Where we’re going
End-to-End Probabilistic Algorithms
Running on GPUs today · FPGA validated · ICML 2026
2D multi-chip SPU
28 nm · custom memory hierarchy
3D-stacked SPU
Near-memory compute · datacenter
Physicists and chip architects, building from first principles.
Flucta's founders started the field of probabilistic computing with p-bits, and many of the field's Nature-level firsts are theirs: the first demonstration of 1,000,000 p-bits, the first carbon nanotube computer, and firsts in 3D chips and in hardware/software co-design.
Founding team
* ConsultingKartik Prabhu
Ph.D., Stanford
Shuvro Chowdhury
Ph.D., Purdue
Corentin Delacour*
Ph.D., Univ. of Montpellier
Kemal Selçuk
Ph.D., UC Santa Barbara
Nihal Sanjay Singh
Ph.D., UC Santa Barbara
Kai Bartolone
M.S., Stanford
Kevin Callahan-Coray
Ph.D. Candidate, UC Santa Barbara
Zengxiao He
M.S., Stanford
Advisors
Funded by investors
who back foundational technology.
Frequently Asked Questions
Flucta builds probabilistic computing across the full stack. Our algorithms and runtime speed up inference end to end on today's GPUs while keeping the model's intelligence intact, and the same principles carry into FPGA prototypes and purpose-built silicon.
Most hardware companies treat algorithms as static benchmarks, while software teams treat silicon as an unchangeable black box. Flucta takes a top-down, full-stack approach. By co-designing software algorithms, runtime SDKs, and physical execution targets around the principles of probabilistic computing, we capture orders-of-magnitude efficiency gains across the entire stack; benefits that isolated software or pure-play chip efforts cannot reach on their own.
Not initially. Though our principles will naturally lead us there in terms of custom ASICs, 3D-CMOS, and near-memory computation, our initial focus operates at the system and architectural level. Instead of waiting on novel physical devices, we re-architect memory hierarchies, routing logic, and execution pipelines on standard silicon processes to natively support sampling and probabilistic operations today.
Enormous efficiency is lost between the design abstractions of the modern stack. Algorithms are written as if moving memory were free, runtimes assume dense and deterministic execution, and silicon is built for worst-case precision. Each boundary hides what the layer below actually needs, and the waste compounds. Flucta removes these boundaries with a top-down design philosophy, carrying one probabilistic abstraction from the algorithm through the runtime into the silicon, so the gains at each layer multiply.
Email founders@flucta.ai to request the technical brief. We are hiring first-principles thinkers across physics, computer science, computer engineering, kernel development, and inference optimization.
