All work

Calx

Shelved

When a repeated correction needs a mechanism.

Calx grew from a question in my own agent-assisted builds: what changes after I correct an agent? I developed a platform for capturing recurring corrections, proposing rules, and enforcing approved behavior in the runtime. The work spans practitioner research, an early CLI, and a later platform prototype.

My role

Founder; product and technical owner

Current state

Calx is largely shelved while I consider whether to open source parts of the work.

Text version
This illustration describes the designed correction lifecycle, from a recorded mistake to a rule enforced by the runtime. Recurrence prompts a suggestion; a person decides whether to promote it.
  1. Capture

    Record a correction with enough context to describe what happened and the behavior that needs to change.

  2. Recognize recurrence

    Compare corrections to identify a repeated issue. Recurrence is a reason to consider a structural change, rather than proof that a rule is ready.

  3. Suggest a rule

    Surface the repeated issue as a candidate for enforcement. The suggestion gives the user a decision to make.

  4. User promotes

    The user chooses whether to promote the correction. Crossing a recurrence threshold alone does not start the compilation process.

  5. Validate

    The compilation path generates and validates an enforcement artifact before distributing an approved rule to the runtime.

  6. Enforce in the runtime

    Middleware applies the rule as the agent works. The application can govern tool availability and actions outside the model's conversational context.

The question behind the product

While building with agents, I kept encountering a frustrating pattern: I would correct a mistake, record the lesson, and see a related mistake later. Instructions and memory were useful, but recording a correction did not tell me whether the system’s behavior had changed.

That became the question behind Calx. When should a correction remain guidance, and when should it become a mechanism the software can enforce?

I owned the work end to end as founder: research direction, product scope, architecture, implementation, and review, using coding agents throughout. The research came from my own builds and their correction records. It was practitioner field research, with the limits of observational evidence from a single operator.

Make the correction lifecycle explicit

The early product was a local CLI that captured corrections and brought derived rules into later sessions. I released that version, then developed a broader platform with a cloud backend, an agent runtime, and a desktop interface.

The later design treats a correction as the beginning of a lifecycle. Capture it, look for recurrence, suggest a rule, let the user choose whether to promote it, then validate and distribute the resulting enforcement artifact. That sequence made several responsibilities explicit that a growing instructions file leaves implicit.

One consequential decision was to keep promotion with the user. A recurring mistake can suggest that something needs to change, but the recurrence threshold should not make that decision on the user’s behalf. The platform surfaces a proposal; the person decides whether it should become enforced behavior.

Own the part that must enforce the promise

The product started around an existing coding-agent environment. As the design developed, the runtime boundary became central. The product needed to govern things such as which tools were available and whether prerequisites had been met before an action could proceed.

I moved toward a dedicated harness so those controls could be implemented in middleware. The desktop client provided the working surface, the harness ran the agent and its controls, and the backend managed the correction lifecycle and rule distribution. That division also kept proprietary compilation work outside the desktop bundle.

This introduced real complexity: multiple runtimes, contracts between them, and packaging requirements. The architectural choice followed the behavior the product needed to support. It also increased the amount of integration work required before the whole system could be considered ready.

A passing test can still describe the wrong system

One of the most useful failures involved an orientation gate intended to withhold mutation tools until the agent had read its briefing. Tests passed, but the implementation relied on a category attribute that the real tool objects did not carry. Some hardcoded tool names also differed from the actual catalog.

The resulting gate could look correct in isolation while doing nothing in the runtime. The fix derived the relevant tools from the shared catalog and tested against real tool objects. The lesson was specific: a control is only as strong as the connection between its test representation and the objects it governs.

What exists, and where it stands

The work produced the released early CLI and a later backend, harness, and local desktop build. The later platform still had integration and verification work remaining at its close-out review.

Calx is now largely shelved. I am considering whether to open source parts of it. The lasting thread is the product question that started it: deciding which behavior can be guided by the model, which should be governed by software, and where the user should retain the decision.