Self-improving coding environment · biofeedback loop
A development environment that reads your heartbeat, learns which interruptions keep you in flow, and evolves with you · under your hand, never on its own.
Where this came from
At Organized AI, at Antler, Esteban reproduced HarnessX live: the idea that you improve an agent by evolving its harness (its tools, memory, prompts, control flow) from real execution traces, running local Ollama models with a frontier model as the judge.
Could the harness evolve from my biofeedback?
Could the environment I build with learn, from my own nervous system, how to keep me in flow? I walked out with that question. This deck is the experiment that answers it.
Esteban’s workshop at Organized AI (Antler), reproducing HarnessX. arXiv:2606.14249. Local model here: qwen3.5:27b.
The invisible tax
A ding arrives. You glance, you answer, you turn back. The switch felt free. It wasn’t.
Studies of interrupted knowledge work put the cost of a single interruption at over twenty minutes to fully return to the task. The body registers the jolt in the first five seconds, long before you notice.
A chest strap can read that score in real time: heart rate, beat-to-beat variability, the shape of the response. The information was always there. Nothing was listening to it.
Two sounds, two heartbeats · real captures
Two notifications, two recordings from the same heart, minutes apart. One kept me in flow. One tore me out of it. My pulse knew which was which before I did.
Heart-rate and HRV values, and the rewards, are the actual measurements from a Polar H10 (RR intervals, single-lead ECG recorded simultaneously). The beat rhythm you see is driven by those real numbers; the waveform is rendered for legibility. Reward is a 0–1 orienting score: higher is calmer.
The orienting response · measured on my own heart
Two real notification captures, minutes apart. Attention without threat decelerates the heart and lifts variability. A startle does the opposite, in seconds, before you notice.
Heart rate slows. Orienting: attention captured, calm held.
Heart rate races. Startle: sympathetic, costly, slow to recover.
Both columns are real measurements from a Polar H10, minutes apart. The reward is the 0 to 1 orienting score the system actually learns from. This is the signal a five-second smartwatch sample cannot see.
Inside the sensor · the capture path
No cloud, no phone app in the middle. The strap speaks straight to the laptop over Bluetooth, and the entire score is computed locally, in the moment.
Close the loop
Every alert is an experiment. The heartbeat is the result. Over time the environment learns which signals you meet with calm.
An honest reward
No rating that session? The reward leans on latency and pulse alone. Absent signals are skipped, never imputed, so the estimate can’t drift toward whatever is most often missing.
A wonderful sound dulls with repetition. Beliefs decay toward the prior, so the system tracks the person you are now, not the one you were last month.
A guaranteed slice of curiosity means no early favorite gets to lock the system in. It keeps asking the question your body keeps answering.
The same three commitments govern how the wider environment scores its own changes. The discipline that keeps a heartbeat honest is the discipline that keeps a self-editing tool honest.
The whole system, one picture
Everything the environment learns converges into a single briefing you read once a week. Each loop measures a different thing; none of them touch the live system on their own.
Loop 3 is deployed today. Loop 1 is built and gated. Loop 2 and the weekly report are the next build. The picture is the destination; the honest status is on a later slide.
You steer the evolution
This is guided evolution, not autopilot. The environment can draft a change to its own instructions, sounds, or workflow, but nothing reaches the live system until it clears three gates · and the last gate is you.
Plain code, not a model. No previously working behavior may regress; a curated safety set must hold with zero tolerance; the change must actually help. Frozen guardrails can never be touched.
A second model from a different family tries to refute the change: is the gain real, or is it gaming the test? It can veto, but it can never overrule the code gate.
Every surviving change arrives as a plain diff with its evidence and a one-line undo. You promote it, or you don’t. Nothing is ever applied on its own.
Not a tradeoff
The old assumption is that productivity is bought with strain. The heartbeat says otherwise: the calm alert is the one you answer fastest, remember longest, and pay for least. Optimize for the regulated nervous system and the other two follow.
Fewer startles, and none during deep focus. The thread stays unbroken.
Higher variability, less sympathetic load, measured not guessed, across the day.
The signal you meet with calm is the one you act on soonest and recover from fastest.
The bigger loop
Notifications were the proving ground. The same architecture · instrument, measure, reward, gate, propose · runs over the environment itself. It reads its own execution traces, digests them locally so nothing private leaves the machine, and drafts small, reviewable improvements to its instructions and skills. Every week it hands you a report: here is what I noticed, here is what I would change, here is the evidence, here is the one command to undo it. You remain the conductor.
The engine that gates changes was itself hardened by that exact method: a model proposed the code, an adversary found sixteen flaws, a test suite gated the fixes, a human merged them. Even this presentation’s companion paper was adversarially reviewed and lost its two best claims for being unprovable. The gates apply to everything, including themselves.
Do this yourself
Observed-only reward. Never impute a signal you did not measure. Missing is missing, not zero.
Recency decay. Let old beliefs fade toward the prior, so the system tracks who you are now.
Propose-only. The machine drafts, a gate filters, an adversary attacks, and a human decides. Every time.
Get those three right and the parts underneath can be anything. The discipline is the product.
The shopping list
Total new spend to replicate the loop: about $90. Everything else is free or software you already run. The strap is the only thing between you and your own biofeedback loop.
Where it is, honestly
Above: an actual single-lead segment recorded during impression 472. Real recording, real heart, real 0.69.
The thesis
Build a development environment that keeps you regulated, keeps you healthy, and gets a little better every week · because it learns from the one signal that can’t lie, and it never changes itself without asking.
SICE · a biofeedback-driven, human-gated self-improving coding environment
the gates are not overhead on the value · the gates are the value
Connect with me
Colin McNamara
Field CTO, AHEAD
Founder & organizer, AIMUG
Founder, Acquit.ai
colinmcnamara.com/connect