# Calx

Turn repeated corrections into runtime controls.

I designed and built an agent workspace around a practical question: what changes after you correct an agent? Calx connects correction capture, deliberate rule promotion, and runtime controls across a desktop interface, a Python harness, and cloud services.

Author: Spencer Hardwick
Role: Founder · End-to-end product and engineering
Status: Development paused
Type: Independent project
Placement: independent

Independent project · development paused. The work spans an early CLI and a later desktop workspace, runtime, and cloud services.

## Workflow explained

Capture a correction, identify recurrence, deliberately promote it, validate the candidate, and distribute a signed rule for the runtime to verify and bind.

### Capture

Record the correction and the context behind it.

### Recognize recurrence

Use related corrections as evidence for a possible rule.

### Human promotion

A person decides whether the correction should become a constraint.

### Validate

Generate and validate a candidate, with review where needed.

### Runtime control

Verify signed rule material locally and bind applicable controls.

## The question behind Calx

I built Calx around a pattern in my own agent-assisted work: a correction could be recorded clearly and still need repeating later. I wanted a product that connected the lesson to a mechanism in the software.

I owned the project end to end: research, product direction, interface design, architecture, implementation, and evaluation. I used coding agents throughout the build, while owning the decisions and reviewing the resulting behavior.

## From instructions to a working surface

The early CLI explored how to capture corrections and carry rules into future work. As the product developed, the interaction needed a home alongside the agent’s actual work. Bench became that surface: a desktop interface for sessions, correction capture, and deliberate promotion.

The architecture evolved with the product. Earlier file, MCP, and hook approaches informed the design, but the later system put execution inside Tether, a Python runtime connected to Bench through a local WebSocket. That gave the product a place to implement controls directly.

## Decide what belongs in software

Some behavior benefits from guidance. Other behavior needs a prerequisite the application can check. Reading a briefing before changing files is a useful example: the runtime can govern tool availability until orientation is complete.

I kept human promotion between correction and compilation. Repetition is useful evidence, but the operator still decides whether the proposed constraint fits the work. That decision is central to the product experience.

## Design the boundary, then test it

The desktop interface, local runtime, and cloud services each have a distinct responsibility. The interface makes the lifecycle understandable. The runtime owns execution. Cloud services handle compilation and signed distribution.

A key engineering lesson was to test controls against the real runtime objects and tool catalog. The meaningful question is whether the mechanism governs the action it was designed to govern. That principle shaped how I approached orientation, tool availability, and action checks.

## What this work represents

Calx brought together product research, a desktop working surface, agent middleware, model routing, rule distribution, and session continuity. Development is paused, and I am considering opening up parts of the work.

The product judgment remains useful across agent systems: decide which behavior needs guidance, which needs an execution boundary, and which decisions should remain with the person using the system.

## Inside Bench

Design prototype · illustrative workflow

One illustrative correction follows the same path through capture, promotion, and intended subsequent behavior: read the project briefing before making changes.

### Capture: Keep the correction beside the work

The operator records that the agent attempted a change before reading the briefing. The correction retains its context in the workspace.

**Before:** Agent attempts a file change before orientation.

**After:** Correction captured: read the briefing before making changes.

### Promotion review: Decide what should become a rule

Related corrections make the issue worth reviewing. The operator deliberately promotes an orientation prerequisite for tools that change state.

**Before:** Related corrections suggest a recurring orientation issue.

**After:** Operator promotes: withhold mutation tools until the briefing is read.

### Subsequent behavior: Make the prerequisite part of execution

The intended interaction gives the agent a clear path: read the briefing, satisfy the orientation gate, then proceed with changes.

**Before:** Mutation tools are withheld while orientation is incomplete.

**After:** Briefing read → orientation satisfied → mutation tools available.

## Architecture

### System architecture

A local working surface and owned agent runtime, with centralized compilation and signed rule distribution.

#### Working surface

**Bench** — Desktop interface

Bench provides the working surface for sessions, corrections, and promotion decisions. The Svelte and TypeScript interface runs inside a Tauri shell.

Technologies: Svelte, TypeScript.

**Tauri shell** — Native bridge

The Rust shell owns the native connection to the local runtime and bridges its events into the interface.

Technologies: Tauri, Rust.

#### Local execution

**Tether** — Agent runtime

The Python runtime owns sessions, middleware, tool availability, model routing, and the local rule cache. This is the execution boundary where application controls meet the agent.

Technologies: Python, LangChain, LangGraph.

#### Services and inference

**Cloud services** — Compilation and distribution

REST services coordinate correction lifecycle work, rule compilation, validation, signing, and distribution. Compilation stays separate from the desktop bundle.

Technologies: Python.

**PostgreSQL** — Lifecycle records

The cloud persists correction and rule lifecycle records. The local runtime maintains the rule material it needs to execute.

Technologies: PostgreSQL.

**Direct inference** — Bring your own key

BYOK routing sends inference from the runtime to the selected model provider.

**Managed inference** — Cloud-proxied route

Managed routing sends inference through the cloud service. That route makes the cloud part of the inference data path.

#### Connections

- **Bench → Tauri shell (Native bridge):** The interface invokes the Tauri shell and receives runtime events through its native bridge.
- **Tauri shell → Tether (Local WebSocket):** A local WebSocket carries commands and events between the Rust shell and Python runtime.
- **Tether → Cloud services (REST + update stream):** REST supports lifecycle operations and artifact retrieval; an update stream announces rule changes.
- **Cloud services → PostgreSQL (Persistent records):** The service stores lifecycle records in PostgreSQL.
- **Tether → Direct inference (Provider API):** BYOK inference routes directly to the model provider.
- **Tether → Managed inference (Managed request):** Managed inference routes through the cloud proxy.

#### Own the execution boundary

I moved beyond instructions and external hooks to a runtime that can govern tool availability and action prerequisites.

#### Separate local execution from compilation

The agent runs locally while centralized services handle rule compilation and signing. That division gives each side a clear responsibility.

#### Make the workspace part of the product

Bench brought the correction loop into the place where work happens, reducing the friction of a separate CLI workflow.

### Correction lifecycle

A repeated correction is evidence for a decision. A person decides when it should become an enforceable constraint.

#### Observe

**Capture** — Correction and context

Record what happened, what should have happened, and the context needed to interpret the correction.

**Recurrence evidence** — Patterns across work

Related corrections provide evidence of a recurring problem. Recurrence informs a suggestion; it does not independently authorize compilation.

#### Decide

**Human promotion** — Deliberate choice

The user chooses whether to promote a correction into a candidate rule. This preserves judgment about scope and consequences.

**Generate and validate** — Candidate artifact

Generate a candidate enforcement artifact and evaluate it. Review is conditional where validation or ambiguity requires a decision.

#### Apply

**Sign and distribute** — Versioned rule material

Approved artifacts are signed and made available to runtimes. A signature establishes artifact integrity, separate from behavioral correctness.

**Verify and bind** — Local runtime

The runtime fetches rule updates, verifies signatures, updates its cache, and binds applicable controls into active sessions.

#### Connections

- **Capture → Recurrence evidence (Related evidence):** Recorded corrections supply the context used to identify recurrence.
- **Recurrence evidence → Human promotion (Suggestion):** A repeated pattern creates a promotion opportunity for the user.
- **Human promotion → Generate and validate (Explicit request):** Only deliberate promotion starts the candidate-generation path.
- **Generate and validate → Sign and distribute (Validated artifact):** Validation informs whether an artifact proceeds or needs review.
- **Sign and distribute → Verify and bind (Update → fetch → verify):** An update event triggers artifact retrieval and local verification before binding.

#### Keep promotion explicit

Repeated mistakes can have different causes. The user decides whether the proposed constraint is the right response.

#### Test behavior as well as integrity

Artifact signing protects provenance and integrity. Validation still has to address what a rule does in the runtime.

### Runtime controls

Put each control at the point where the runtime can act on it, and carry useful context across sessions.

#### Before action

**Orientation** — Establish context

An orientation gate checks prerequisites such as reading a briefing before exposing tools that can change state.

**Tool availability** — Available choices

Middleware controls which tools are offered to the agent, using the runtime tool catalog rather than relying only on conversational instructions.

#### At execution

**Action checks** — Execution boundary

Check applicable conditions when an action is attempted. This makes the runtime responsible for enforcing the boundary.

#### After action

**Response review** — Output boundary

Response review provides a separate place to evaluate the agent’s output against applicable constraints.

**Structured handoff** — Session rollover

A structured handoff carries relevant working context into a subsequent session. Session rollover is an explicit part of the harness lifecycle.

#### Connections

- **Orientation → Tool availability (Prerequisites satisfied):** Orientation status informs tool availability.
- **Tool availability → Action checks (Selected tool call):** An available tool can be selected, then checked at the execution boundary.
- **Action checks → Response review (Execution result):** Results return to the agent and pass through the response boundary.
- **Response review → Structured handoff (Carry context forward):** The session can preserve a structured account of the work for continuity.

#### Match the test to the runtime

Controls need tests against the actual tool catalog and runtime objects. A convincing isolated test is useful only when it represents the execution boundary.

#### Treat continuity as a system responsibility

Long-running work needs an intentional handoff between sessions, with useful context preserved in a structured form.



---

[View this case study](https://spencerhardwick.com/work/calx/)
[All work](https://spencerhardwick.com)
