Interface Shapes Cognition
A reinforcement-learning experiment shows why the environment carries the five in the 1:3:5 Rule — and why output problems are usually container problems.
In the Conscious Stack, the 1:3:5 Rule holds that a working system has one Anchor, three Active moves, and five Supporting elements. There's a working rule underneath it that doesn't get said often enough: memory carries the 1:3, and the environment carries the 5. The five aren't things you remember. They're installed in the space around you — in defaults, in layouts, in the shape of your tools — precisely so you never have to recall them. The Anchor and the three moves are the parts you can recite at 11pm without opening the page. The rest has to live somewhere that isn't your head.
Most of the argument for this has come from lived evidence — stack audits, watching capable people drown in well-chosen tools arranged badly. But in late 2024, a small reinforcement-learning experiment produced the same shape with numbers attached, and it's worth sitting with.
The experiment
In November 2024, Kim and Sung at Sangmyung University published a paper in Information on a niche problem: teaching game agents to navigate hexagonal maps in Unity's ML-Agents toolkit. Hex grids are everywhere in strategy games, and the toolkit had no good built-in way to sense them. So they ran three agents on the same goal-reaching task, with the same learning algorithm, and changed only one thing — how the map was described to the agent.
Agent A got the standard approach: a buffer sensor with axial coordinates. Final reward: 0.506. Agent B got a better sensor — a purpose-built hexagon sensor — but kept the coordinate-based description. It scored 0.186. Worse, with better hardware, so to speak. Agent C got the exact same sensor as B, but the map was described in concentric layers radiating outward from the center, indexed by layer instead of by coordinates. It scored 0.994, reached the goal roughly eight times faster than A and fifteen times faster than B, and trained in about half the time of the standard setup.
The word the paper uses for what made the difference is "alignment." Not AI-safety alignment — nothing here is about values, or goals, or intent. It's index alignment: every cell of the map occupies the same slot in the list the agent reads, every single time, and its position can be inferred from that slot alone. The raw version of this is a seating chart where all the names have been shifted one seat over. The person reading it isn't stupid. The chart is lying to them, politely, and they walk to the wrong chair with full confidence.
The turn
The pair that matters is B versus C. Identical sensor. Identical task. The only thing that changed was the geometry of the description — the container the information arrived in. C's description pre-loaded position into the shape of the data itself, so the agent never had to spend capacity figuring out where it was. It could spend everything on what to do next.
That's the five, doing their job. When a system is index-aligned — when the thing is where the structure says it should be — cognition stops paying tax on orientation and starts spending on the actual move. You don't feel the support when it's working. You just think clearly, and it feels like you. When the alignment is off, everything reads as a discipline problem. You're slower, you second-guess, you burn attention reconstructing context that should have been carried by the environment. It was never a discipline problem. It's a container problem.
One thing worth noting: this is a single experiment, on one navigation task, in one game engine. It doesn't prove the 1:3:5 geometry, and the Rule doesn't rest on it. The framework is informed by bounded cognitive capacity, hierarchical attention, and patterns observed through stack audits. What's interesting here is narrower and more useful: a system simple enough to measure, producing exactly the shape the 1:3:5 Rule predicts.
The rule
When output degrades, audit the interface before the agent. The container before the discipline. The chart, not the person.
Most people — and most teams — run this in reverse: they retrain, restate goals, add willpower, and leave the scrambled observation layer untouched. The hexagon experiment is the cleanest small demonstration I've seen of why that order fails. Same learner, same goal, same sensor. Change the shape of what it's shown, and you change what it's capable of learning.
Environmental geometry holds certain memory effortlessly, for those in it or passing through it.
Want to see the shape of your own stack? Map it in the Stack Builder, then look at what you've placed where.