Φ Institute for Physical AI @ JBI · Charlot Lab

Energy First Architecture

One energy a body descends to act, and the same energy is the proof it won't diverge.

The dominant AI stack (point neuron, dense matmul, frozen weights, a black box in a cloud datacenter) was built for machines that talk. Physical AI is for machines that act, and a wrong action has physical consequences. The incumbent commit rule is an expected-value score, never a guarantee, never before the action commits, never on the robot. EFA is the on-device verification layer that closes exactly that gap: the same scalar energy a body descends to act is a machine-checked proof the closed loop will not diverge, in microjoules, on the device, before the action commits. One object is both the controller and its own certificate; joules per task is the cost of carrying the proof, not a throughput contest.

certificate · a machine-checked proof the action won't diverge, before commit
controller · the same energy, descended to a goal, is the policy
world model · predict consequences by descending it

The thesis

One energy, both roles, the controller is its own certificate

The claim is unification, and it is narrow on purpose. One scalar energy over a sparse-positive latent is at once the policy a body descends to act, the model it predicts consequences in, and, and, the part that distinguishes it, a machine-checked certificate that the action won't diverge before it commits. The proof is not bolted on; it is the objective. The on-device, sparse-compute substrate that lets this run in a browser tab is a separate adjacent, energy-native compute, not part of this certificate claim.

layerincumbentEFA
guaranteeexpected-value score, after the factthe objective is the certificate, a proof before the action commits
world modelautoregressive next-tokenpredict consequences by energy descent, the same energy
unitpoint neuronsparse-positive, monosemantic
memoryfrozen weights + finite contextHebbian fast-weights, written at inference
legibilityblack boxmonosemantic / steerable by construction

What's proven, measured, and priced

The edge is narrow and real

Built the way it should be judged: prove each mechanism, compose, scale, push to hard tasks and real data, pricing every limit. EFA is not claimed to beat frontier systems at scale. Its demonstrated edge is on the axes where a native energy is the right representation.

axisrepresentative result
the true Energy-Based Transformertrained through an unrolled energy descent (2nd-order autograd): solves a multivalued system by representing its whole solution set, and thinks, accuracy climbs with compute (K=1→6: 22→100%)
test-time thinking (acting)MPPI latent planning 39% → 69% reach, same value net
native verificationenergy best-of-N → 100% selection; on real AR sequences too
zero-shot compositionconcept conjunctions never trained jointly, by energy summation
OOD generalizationgoal-agnostic energy 37→41% OOD where learned maps collapse
energy conservationHamiltonian NN, 5.5× lower energy drift
scientific discoveryLorenz · Burgers · Fisher–KPP · real lynx–hare data, laws recovered
intelligence = economy of effortas the model learns, the compute to solve falls (descent steps 8→4), the landscape recodes so the hard problem becomes easy: fewer watts, not a bigger model
materiality does computationphysical descent packs discs 100% where abstract search collapses to 0%, the body drops the search clauses (Krakauer's Soma cube)

Grounded in the complexity-science account of intelligence (Krakauer, SFI): the physics → adaptation → agency least-action ladder made computational; intelligence as making hard problems easy. The edge is a better-shaped energy, not a bigger model.

Where the field is, and where EFA fits

The field found energy-as-score. EFA also makes it the certificate.

The convergence is real, and we build on it. A softmax is the Boltzmann distribution, take the energy to be the negative logit, so every classifier is already an energy model. JEPA makes the energy a distance in latent space and kills the intractable partition function Z. An energy-based transformer thinks by descending its own energy at inference. Concepts compose by adding energies. All of it is now mainstream, and all of it scores: it rates compatibility, density, plausibility. EFA does that too (the table above), but that is table stakes, not the claim.

The claim is the connection the scorer wave never makes: the same energy a body descends to act is also a machine-checked certificate that the action will not diverge, before it commits. The scorer rates an output after the fact; the certificate is a guarantee, in the loop, on the robot. Same word “verify,” different object:

“verify”energy-as-score, the fieldenergy-as-certificate, EFA
asksis this output a good answer?will this action keep the closed loop from diverging?
isa score on a guess, density, distance, compatibilitya Lyapunov guarantee, V̇ ≤ 0 by construction
holdsafter the answer is producedbefore the action commits, at the real dt, on the device

Energy-as-score is crowded and real; we do it, and we don't lead with it. Energy-as-certificate for embodied control is an active field of its own: learned control-barrier and Lyapunov certificates, runtime safety monitors for VLAs, and port-Hamiltonian energy certificates are all being built right now. EFA does not claim that ground. Its contribution is to make it open and on-device: the same energy serving as the controller and its own proof, in pure Rust you can run in a browser and read end to end.

The necessity: a score is not a proof

The commit rule is a score. A body needs a proof.

A robot's wrong action has physical consequences. The incumbent stack (predict, score the rollouts, commit the best-looking one) offers only an expected-value score: never a guarantee, never before the action commits, never on the robot (LeCun's "intrinsically unsafe"). Closing that gap is an active 2025–26 effort in its own right: learned control-barrier certificates, runtime safety monitors for VLAs (SafeVLA, Pre-VLA, AEGIS), and port-Hamiltonian energy certificates. EFA's answer is the open, on-device, physics-first one: the same energy a body descends to act is a machine-checked proof the closed loop will not diverge, in microjoules, on the device, and able to wrap a frozen policy it did not train. No bolted-on verifier: one energy, and the proof is the objective.

the claimproven
structure, not proofstructure the energy as a port-Hamiltonian and the Lyapunov certificate holds by construction, V̇ ≤ 0 to machine precision at a 24-dimensional state, where a box/SMT proof dies at ~8D. And it survives a learned residual: learn what you don't know, keep the guarantee.
necessity, not a weak baselineon a 7-DOF arm, an aggressive task policy leaves the safe envelope 6126 / 12000 steps; the energy certificate vetoes the unsafe actions and holds it at 0 / 12000, a guarantee before commit, 143 µs/step, on-device.
on a real robot bodythe same gate runs on a MuJoCo arm, real inertia, motor armature, torque-limited motors, the energy built from MuJoCo's own mass matrix. The structural identity holds to 1.5e-9 on the real dynamics; a naive policy leaves the envelope 4263 / 6000 steps, the certificate holds it 0 / 6000, zero unrecoverable, at the real dt.
how small the seatbelt can getthe certificate is claimed to be light, so we measured how far it compresses before the guarantee goes. A learned energy behaves like any other network and worse: squeezing the trained weights after the fact drops certified coverage to a median in the low sixties, essentially the 61% the trivial V=|x|² already scores, truncation does not degrade the learned certificate, it destroys it. Training the same budget in factored form from the start holds 100% of the domain at every rank, across all five seeds, at 16× fewer parameters. A structural certificate has no such failure: place a learned residual in the skew coupling and it cancels out of V̇ identically, worst V̇ is −2.8 whether the residual keeps 64 units or 1, while wiring the same residual into the damping breaks the guarantee at every compression level. The rule: compress freely whatever the guarantee does not depend on; anything it does depend on must be trained small, never shrunk afterwards.
the observation it was handedevery certificate above assumes the state it receives is true. A body in the field is handed an estimate, and sensors degrade and get spoofed, so a certificate on a corrupted state is not noisy, it is confidently wrong, reporting safe while the body is in violation. Measured over 60,000 trials: under scattered sensor faults a naive barrier reaches 16% false-safe, a trust-weighted energy fuser plus a plausible-set margin drives that to ~0% for a deliberate ~20% conservatism. Under a coordinated spoof the comfortable result breaks, ours included: when one transmitter captures five of nine sensors they agree with each other, the fuser locks onto the self-consistent lie, and trust-gating scores worse than the naive mean it was meant to fix (36% vs 28% false-safe). What survives is refusal, detecting a rival, internally-consistent explanation and declining to certify cuts false-safe to under 2%, at 95% abstention. Past a coordinated majority the sensors look like noise, the detector goes blind, and everything sits near 34%: recoverable only by an anchor the adversary cannot forge. Refusal is not a failure mode of the certificate; it is part of it. Watch it fail live →

Several 2025–26 lines are converging on exactly this, including the port-Hamiltonian energy-certificate work below. EFA is our open, on-device take on that convergence: one energy, the proof structural, the whole loop runnable and readable in a browser. It builds on these; it does not claim to precede them:

EBT-Policy · energy-equilibrium control beats diffusion, but the energy is a scorer, not a Lyapunov certificate (2510.27545)
ECD · long-horizon = a sum of local energy potentials, convergent with EFA; we extend, not beat (2606.21646)
V-JEPA 2 · a real-robot planner that minimizes a latent energy, a proto-EFA, but frozen and pixel-free, no certificate
Certified World Models · a Lyapunov certificate on the model, but certificate and action-objective stay two separate scalars (2606.13092)
Port-Hamiltonian certificates · co-learned pH model + energy-shaping control, provably passive, keeping the natural potential (2604.26172, 2512.24493), the structural line EFA implements openly and on-device
Runtime safety for VLAs · preemptive verification and backup reflexes around foundation-model policies (SafeVLA 2503.03480, Pre-VLA 2605.22446, AEGIS 2606.06660), the verification field EFA teaches, not owns
Certificates under untrusted observation · measurement-robust and fault-tolerant barrier certificates, secure safety filters under sensor spoofing, and perception-based safety under out-of-distribution measurements (2402.18677, 2505.06845, 2511.08741, 2602.05311), a mature line EFA claims no primacy over; our contribution is the open, readable, on-device version, measured with its failures reported

Scope, drawn correctly. The certificate is lightweight by design (microjoules, O(1)), and that is the feature: a small, fast proof that wraps a full-scale policy it did not train (a released VLA of any size) and holds at any number of joints for a fully-actuated body, by an exact identity rather than a size-limited approximation, the structural (port-Hamiltonian) certificate is dimension-free (verified past a 24-dim state; the algebra does not stop there). What is bounded is specific and priced: the exhaustive worst-case proof of a learned energy is 2–4D (exactly why we move to structure, not proof); underactuated bodies re-introduce the hard part; results are in simulation, not yet on hardware; contact-force / manipulation-task certification is the open frontier; and the observation layer has a measured hard edge, beyond a coordinated majority of spoofed sensors no pipeline of ours recovers the state, because the readings no longer contain it, so that regime needs an unforgeable anchor (different-physics sensing, attestation, a hard dynamics prior) rather than better inference. The gate already runs on a real MuJoCo body against a policy stand-in; the one increment left in the decider is a released external VLA in the same loop (weights on disk; the gate reads state + a proposed action, so it is runtime integration, not a proof change). The limit is never that the AI must be small, the AI it certifies is as large as you like; the seatbelt is what stays light.

Living papers, start with The Certificate Doctrine (the whole thing in one place): Energy Is the Certificate · One Energy, Both Roles · The Grasp Certificate · The No-Harm Certificate · An Entity, Not a Controller · What Directed Curiosity Is For · The Agency Ladder · drive the certificate-gated imagination sim, or watch The Certificate, Live (a fleet that never collides).

Architectural match, the embodied loop

Energy is the right shape for a body

The deeper reason a cloud LLM fails in a robot is architectural: a token model has no energy to shape, no hidden state to infer, no dynamics to descend. EFA's loop is one energy object, perceive, model, and control a real underactuated body, on-device. Each row below is measured; each caveat is priced in the ledger.

capabilitymeasured result
control by energyunderactuated pendulum swing-up from a learned energy; a position controller with the same actuator can't reach upright, it has no energy to pump
perception = inferencecart-pole balanced from noisy position-only obs by inferring the state (an energy observer); naive velocity-differencing falls over at the same noise
full regulationthe whole underactuated cart-pole, balance and cart position, regulated from noisy positions, via a proper LQR
energy of a chaotic bodythe conserved energy of a chaotic double pendulum recovered from trajectories alone (corr 0.998), never told the formula

Performance per watt, quantified. A whole pendulum swing-up under the energy controller costs ≈110 kFLOP, microcontroller scale, against ≈14 GFLOP for a single token of a 7B model: roughly 10⁵× less compute, and the cloud model can't run the continuous loop at all. That is the guarantee being cheap by design, a proof light enough to run on the robot, in the loop, beside a policy of any size. The point is not that the controlled task is small; it is that the seatbelt is.

Adjacent, the same energy, turned on data

Energy-minimization is also law discovery

A separate use of the same machinery, not the certificate claim above, and we keep it distinct on purpose. Fit the energy to data instead of a body and its minimum is a governing law. Honest about the boundary: this is an adjacent capability, not what EFA owns.

Discover the equation

From noisy data, recover the governing ODE, a 2D oscillator, and the Lorenz chaotic system, exactly.

Discover the PDE

Burgers' (advection–diffusion) and Fisher–KPP (reaction–diffusion), recovered exactly from spatiotemporal fields.

Discover the invariant

A conservation law for the nonlinear pendulum, its energy, found from trajectories, correlation 0.99, without its form.

Discover from real data

Predator–prey Lotka–Volterra dynamics recovered from the real Hudson Bay lynx–hare records (1900–1920).

Relation to IPAI @ JBI and Physical AI

Why the Institute builds this

The target is total cost of operation, not a benchmark score. The question EFA is built to answer is what a task costs to perform — joules on the device and capital in the machine that runs it — and the goal is a stack where whole classes of today's work simply stop being necessary because a cheaper physical route exists. That is a measurable target rather than a slogan: energy per completed task, and the acquisition cost of the hardware that completes it. The Institute teaches the metric first because it is the one a student can carry to any architecture, including ones that do not exist yet.

Incumbent hardware is an artifact of investment, not a proof of optimality. The stored-program machine with a single path between memory and arithmetic is a design decision recorded in von Neumann’s 1945 EDVAC draft; Backus named its cost the “von Neumann bottleneck” in his 1977 Turing lecture and argued the programming model had been shaped by it ever since. Dennard’s 1974 scaling rules made that architecture cheaper every generation without changing it, and when Dennard scaling broke down in the mid-2000s the industry responded by adding cores rather than revisiting the design. x86 became the default through the IBM PC and the manufacturing capital behind it, and CUDA became the default for parallel numerics through a decade of ecosystem investment. Each of these won for reasons an economist recognises — installed base, tooling, capital already sunk — and none of the reasons is a physical argument that the arrangement is the cheapest way to compute.

The physical headroom is measurable. Landauer’s bound puts the thermodynamic floor for erasing one bit at kBT ln 2, about 2.9 × 10−21 J at room temperature. Contemporary digital arithmetic spends on the order of picojoules per operation, five to six orders of magnitude above that floor, and most of it goes to moving operands rather than to the arithmetic itself. A gap of that size is not a rounding error in an efficiency curve; it is the space in which a different arrangement of matter can win. Whether any particular architecture claims that space is an empirical question, which is why this page reports measurements rather than projections.

On-device, everywhere. Physical AI must act at the edge, at low energy, the certificate is only worth something if it runs on the robot. EFA runs in a browser tab on WebGPU via the Institute's pure-Rust Ferric compute layer, on-device, no cloud. (The compute substrate itself is a separate adjacent, energy-native compute.) The newest physics-native substrate, thermodynamic sampling, has a live instrument on the energy program: run a verified Gibbs sampler on your own device, priced in joules, with every claim labelled by its provenance.

Energy is the native language of the physical world. Robotics, control, and the physical sciences are already energy-based. EFA doesn't translate physics into a loss; it represents it as an energy, which is why the discovery suite works cleanly, the prior native rather than bolted on.

Legible, and teachable. Every mechanism is a small, driveable, on-device artifact. EFA is a research bet and a curriculum surface at once, the Institute teaches the ideas by letting students run them.

Where the evidence ends

Where the work stands

Corrected under test. An earlier version of this record asserted a task where "feedforward scores 0%." measured, a plain feedforward solves it, the claim was wrong and has been withdrawn. The energy architecture's real edge is representing multivalued solution sets, thinking at inference, and above all matching the physics of a body, not out-scoring feedforward on a discrete puzzle. Reporting the walk-back is the point.

Chains of coupled constraints. A chain of D coupled links solves only when every link holds at once, so if the links fail independently the whole-chain rate is the per-link rate raised to the D−1 power. Measured across every configuration we ran, the whole-chain score sits at or just above that value, which makes it a usable floor: to hold 90% over six links needs 97.9% per link, and over twelve links 99.0%. That single relation explains a result that otherwise looks arbitrary. Widening the network lifted per-link accuracy fivefold at a six-link chain and moved the whole-chain score not at all, because pD−1 punishes chain length faster than width buys accuracy. Improving the link instead took the same chain from 2% to 96%, and a sixteen-link chain to 57%, with every row landing on the floor its own per-link rate predicts.

The link's difficulty turned out to be governed by how many answers satisfy it rather than by the chain's length. Holding dimension, range and smoothness fixed and only multiplying the number of admissible solutions, per-link accuracy holds near 96–97% at two and four solutions and collapses into the low teens at twelve, while an oracle confirms a solution exists in every case. Two abilities separate at that point: the same energy ranks a true pair above a random invalid one 97% of the time while its global minimum lands on a valid point 12% of the time. Ranking is a local pairwise comparison; placing the minimum is a global statement about the landscape, and for a model whose output is defined by minimisation only the second one counts.

The cause is representational, not procedural. A smooth network reading raw coordinates has to synthesise the admissible set’s angular structure out of Cartesian inputs; a sinusoidal network of the kind used for images and audio recovers roughly 40–46% of the gap from the same inputs, and supplying the angular basis directly solves the task completely with a tenth of the parameters of the widest raw-coordinate model. Capacity was never the binding constraint — it becomes useful only once the representation can express the answer, which is why width helped the sinusoidal networks and did nothing monotone for the smooth ones. Figures are medians over four seeds; the low-accuracy comparisons were replicated because seed variance is large there, the high-accuracy chain rows are single runs where it is not.

Sampling. Energy descent (generation) is step-size sensitive; the robust route is verification (best-of-N) or Metropolis correction, priced, not hidden.

A settled negative. An energy-conserving surrogate at field scale did not beat a naive force net, and it still lost after Ferric gained second-order autograd and the test was redone with exact gradients. So the negative is structural, not a tooling gap. (That same autograd unlocked the true EBT above.)

AI-for-math. The energy verifier is mechanism-sound but gated on a Lean-task-trained encoder; general embeddings sit at chance, because tactic↔goal compatibility is a formal, not surface-semantic, property.

Energy First Architecture · Charlot Lab · Institute for Physical AI @ John Bailey Institute.
Whitepaper · Validation ledger · Repository · Open weights, live · efa-hybrid-arm2 · efa-flow-arm3 · Live research record · Certificate-gated imagination · The certificate under spoofing · Why long chains break · Ferric · Institute
Every figure here is a measured result, not an estimate.