Calx
Shelved
When a repeated correction needs a mechanism.
Calx grew from a question in my own agent-assisted builds: what changes after I correct an agent? I developed a platform for capturing recurring corrections, proposing rules, and enforcing approved behavior in the runtime. The work spans practitioner research, an early CLI, and a later platform prototype.
My role
Founder; product and technical owner
Current state
Calx is largely shelved while I consider whether to open source parts of the work.
Select a stage to explore the diagram.
Capture
Record a correction with enough context to describe what happened and the behavior that needs to change.
Recognize recurrence
Compare corrections to identify a repeated issue. Recurrence is a reason to consider a structural change, rather than proof that a rule is ready.
Suggest a rule
Surface the repeated issue as a candidate for enforcement. The suggestion gives the user a decision to make.
User promotes
The user chooses whether to promote the correction. Crossing a recurrence threshold alone does not start the compilation process.
Validate
The compilation path generates and validates an enforcement artifact before distributing an approved rule to the runtime.
Enforce in the runtime
Middleware applies the rule as the agent works. The application can govern tool availability and actions outside the model's conversational context.
The question behind the product
While building with agents, I kept encountering a frustrating pattern: I would correct a mistake, record the lesson, and see a related mistake later. Instructions and memory were useful, but recording a correction did not tell me whether the system’s behavior had changed.
That became the question behind Calx. When should a correction remain guidance, and when should it become a mechanism the software can enforce?
I owned the work end to end as founder: research direction, product scope, architecture, implementation, and review, using coding agents throughout. The research came from my own builds and their correction records. It was practitioner field research, with the limits of observational evidence from a single operator.
Make the correction lifecycle explicit
The early product was a local CLI that captured corrections and brought derived rules into later sessions. I released that version, then developed a broader platform with a cloud backend, an agent runtime, and a desktop interface.
The later design treats a correction as the beginning of a lifecycle. Capture it, look for recurrence, suggest a rule, let the user choose whether to promote it, then validate and distribute the resulting enforcement artifact. That sequence made several responsibilities explicit that a growing instructions file leaves implicit.
One consequential decision was to keep promotion with the user. A recurring mistake can suggest that something needs to change, but the recurrence threshold should not make that decision on the user’s behalf. The platform surfaces a proposal; the person decides whether it should become enforced behavior.
Own the part that must enforce the promise
The product started around an existing coding-agent environment. As the design developed, the runtime boundary became central. The product needed to govern things such as which tools were available and whether prerequisites had been met before an action could proceed.
I moved toward a dedicated harness so those controls could be implemented in middleware. The desktop client provided the working surface, the harness ran the agent and its controls, and the backend managed the correction lifecycle and rule distribution. That division also kept proprietary compilation work outside the desktop bundle.
This introduced real complexity: multiple runtimes, contracts between them, and packaging requirements. The architectural choice followed the behavior the product needed to support. It also increased the amount of integration work required before the whole system could be considered ready.
A passing test can still describe the wrong system
One of the most useful failures involved an orientation gate intended to withhold mutation tools until the agent had read its briefing. Tests passed, but the implementation relied on a category attribute that the real tool objects did not carry. Some hardcoded tool names also differed from the actual catalog.
The resulting gate could look correct in isolation while doing nothing in the runtime. The fix derived the relevant tools from the shared catalog and tested against real tool objects. The lesson was specific: a control is only as strong as the connection between its test representation and the objects it governs.
What exists, and where it stands
The work produced the released early CLI and a later backend, harness, and local desktop build. The later platform still had integration and verification work remaining at its close-out review.
Calx is now largely shelved. I am considering whether to open source parts of it. The lasting thread is the product question that started it: deciding which behavior can be guided by the model, which should be governed by software, and where the user should retain the decision.