Skip to content
Return to Mission Control

Selected system / personal and professional work

Built and used

Harness engineering for AI-assisted software development

I build and use a private engineering environment with deterministic repository intelligence, bounded agent work, explicit orchestration, validation, independent review, and evidence capture.

The problem

An open-ended model request can spend inference on repository details that software can discover deterministically, while hiding dependencies, validation, review, and failure states. Engineering work benefits when those steps are visible and governed by the surrounding system.

Scope

Much of this environment is in daily use. Nearly all AI-assisted development runs in a dedicated server-side sandbox with isolated execution, controlled access to repositories, tools, and credentials, and local or hosted model routing. Deterministic discovery, focused context, orchestration, validation, review, and evidence capture make runs repeatable while I change tools without destabilizing my workstation.

My contribution

I treat agents as bounded components inside an engineered workflow. The system makes discovery, delegation, checks, review, and evidence explicit instead of relying on an unbounded autonomous worker.

Architecture and technical decisions

  1. Use deterministic code analysis, repository graphs, and retrieval to discover relevant structure before asking a model to reason over it.

  2. Construct focused context for each task so model calls receive the evidence needed for that step rather than an entire repository by default.

  3. Use explicit orchestration for dependencies, retries, validation gates, human approval, and failure handling while allowing agents to perform bounded reasoning and tool use.

  4. Separate candidate generation, deterministic validation, and independent review so each produces evidence that can be inspected.

  5. Escalate ambiguous or consequential work to a person and preserve execution evidence for later evaluation.

Constraints and qualifications

  • Agents receive bounded assignments and defined capabilities rather than open-ended ownership of work.
  • Nearly all AI-assisted development runs in a dedicated server-side sandbox with controlled repository, tool, and credential access; I can change tooling without destabilizing my workstation.
  • Discovery, validation, independent review, and human escalation remain visible parts of the workflow.
  • Specialist-model training is a future experiment; first I need enough validated execution evidence.

Guiding principles

  • Use a small deterministic tool first when it can do the work; bring in frontier reasoning when the task needs it.
  • Use explicit orchestration to make dependencies, retries, validation, and failure states visible.
  • Capture each execution as evidence that can support future evaluation and routing.

System views

The model is one component in an engineered system.

I build and use this harness daily in my personal engineering workflow: a dedicated server-side sandbox, deterministic discovery, scoped context, model choice, bounded agents, validation, review, and evidence.

Harness engineering — conceptual system topologyAn architectural illustration, not live telemetry: deterministic discovery narrows the task before model work, then orchestration carries it through tools, validation, review, and evidence. The specialist-model node shows a future experimental path. Execution mode: continuous. Components: Repository intelligence (Code graph · indexes); Deterministic discovery (Find relevant structure); Task-scoped context (Evidence · retrieval); Model routing (Capability · cost · risk); Local model (Private · low latency); Specialist model (Future · after evaluation); Frontier model (Harder reasoning); Bounded agent (Assigned work · tools); Defined capabilities (Scoped tool access); Workflow orchestration (Dependencies · retries); Deterministic validation (Tests · contracts · checks); Independent review (Inspect evidence); Human escalation (Resolve risk or ambiguity); Execution evidence (Context · actions · outcome); Evaluation and learning (Routing · future datasets). Ordered stages: Repository intelligence; Deterministic discovery; Task-scoped context; Model routing; Local model and Frontier model; Bounded agent; Defined capabilities; Workflow orchestration; Deterministic validation; Independent review; Execution evidence. Feedback loop: Execution evidence; Evaluation and learning; the next continuous pass explicitly revisits Task-scoped context; Model routing; Local model and Frontier model; Bounded agent; Defined capabilities; Workflow orchestration; Deterministic validation; Independent review; Execution evidence. Connections: Repository intelligence to Deterministic discovery; Deterministic discovery to Task-scoped context; Task-scoped context to Model routing; Model routing to Local model; Model routing to Specialist model (conditional); Model routing to Frontier model; Local model to Bounded agent; Specialist model to Bounded agent (conditional); Frontier model to Bounded agent; Bounded agent to Defined capabilities; Defined capabilities to Workflow orchestration; Workflow orchestration to Deterministic validation; Deterministic validation to Independent review; Independent review to Human escalation (when required) (conditional); Independent review to Execution evidence; Execution evidence to Evaluation and learning (feedback); Evaluation and learning to Model routing (feedback); Evaluation and learning to Task-scoped context (feedback); Human escalation to Execution evidence (accepted decision) (conditional). Optional conditional paths are shown but do not run in the illustrated sequence.RepositoryintelligenceCode graph · indexesDeterministicdiscoveryFind relevant structureTask-scoped contextEvidence · retrievalModel routingCapability · cost · riskLocal modelPrivate · low latencySpecialist modelFuture · after evaluationFrontier modelHarder reasoningBounded agentAssigned work · toolsDefined capabilitiesScoped tool accessWorkfloworchestrationDependencies · retriesDeterministicvalidationTests · contracts · checksIndependent reviewInspect evidenceHuman escalationResolve risk or ambiguityExecution evidenceContext · actions ·outcomeEvaluation andlearningRouting · future datasets
  • idle
  • incoming
  • active
  • outgoing
  • settled

Conceptual sequence · not live telemetry.

Execution sequence

  1. 01
    Repository intelligenceRepository intelligence → Deterministic discovery
  2. 02
    Deterministic discoveryDeterministic discovery → Task-scoped context
  3. 03
    Task-scoped contextTask-scoped context → Model routing
  4. 04
    Model routingModel routing → Local model · Model routing → Frontier model
  5. 05
    Local model + Frontier modelLocal model → Bounded agent · Frontier model → Bounded agent
  6. 06
    Bounded agentBounded agent → Defined capabilities
  7. 07
    Defined capabilitiesDefined capabilities → Workflow orchestration
  8. 08
    Workflow orchestrationWorkflow orchestration → Deterministic validation
  9. 09
    Deterministic validationDeterministic validation → Independent review
  10. 10
    Independent reviewIndependent review → Execution evidence
  11. 11
    Execution evidence

Ongoing learning loop

  1. 01
    Execution evidenceExecution evidence → Evaluation and learning
  2. 02
    Evaluation and learningEvaluation and learning → Task-scoped context · Evaluation and learning → Model routing

The next pass explicitly revisits Task-scoped context → Model routing → Local model + Frontier model → Bounded agent → Defined capabilities → Workflow orchestration → Deterministic validation → Independent review → Execution evidence for another pass.

Conditional paths

  • Model routing → Specialist model · Future · after evaluation
  • Specialist model → Bounded agent
  • Independent review → Human escalation · Resolve risk or ambiguity · when required
  • Human escalation → Execution evidence · accepted decision
Conceptual execution · explicit dependencies · not live telemetry

Narrow the problem before spending model reasoning.

I combine deterministic code discovery, repository graphs, indexes, and retrieval to assemble task-scoped evidence before a model reasons.

Context and retrieval — conceptual system topologyAn architectural illustration, not live telemetry. Context sources retain provenance and task scope; application services remain authoritative for domain data and rules. Execution mode: once. Components: Task and question (Scope · acceptance); Deterministic discovery (Structure · symbols); Code graph and indexes (Relationships · location); RAG and retrieval (Relevant source evidence); Structured data (Authoritative service); Vector search (Semantic candidate recall); Provenance checks (Source · recency · access); Task-scoped context (Enough evidence to reason); Model reasoning (Focused input · lower waste). Ordered stages: Task and question and Structured data and Vector search; Deterministic discovery and Provenance checks; Code graph and indexes; RAG and retrieval; Task-scoped context; Model reasoning. Connections: Task and question to Deterministic discovery; Deterministic discovery to Code graph and indexes; Code graph and indexes to RAG and retrieval; Structured data to Provenance checks; Vector search to Provenance checks; RAG and retrieval to Task-scoped context; Provenance checks to Task-scoped context; Task-scoped context to Model reasoning. Optional conditional paths are shown but do not run in the illustrated sequence.Task and questionScope · acceptanceDeterministicdiscoveryStructure · symbolsCode graph andindexesRelationships · locationRAG and retrievalRelevant source evidenceStructured dataAuthoritative serviceVector searchSemantic candidate recallProvenance checksSource · recency · accessTask-scoped contextEnough evidence to reasonModel reasoningFocused input · lowerwaste
  • idle
  • incoming
  • active
  • outgoing
  • settled

Conceptual sequence · not live telemetry.

Execution sequence

  1. 01
    Task and question + Structured data + Vector searchTask and question → Deterministic discovery · Structured data → Provenance checks · Vector search → Provenance checks
  2. 02
    Deterministic discovery + Provenance checksDeterministic discovery → Code graph and indexes · Provenance checks → Task-scoped context
  3. 03
    Code graph and indexesCode graph and indexes → RAG and retrieval
  4. 04
    RAG and retrievalRAG and retrieval → Task-scoped context
  5. 05
    Task-scoped contextTask-scoped context → Model reasoning
  6. 06
    Model reasoning
Conceptual execution · explicit dependencies · not live telemetry

Intended effect

The system is designed to make AI-assisted engineering more bounded, reviewable, and improvable. Agent execution should leave evidence that supports reliability analysis and future workflow improvements instead of disappearing as disposable logs.

Technologies and components

  • Repository graphs
  • Workflow orchestration
  • Local inference
  • Independent review