Skip to content
Return to Mission Control

Selected system / personal and professional work

Experiment in progress

Specialized local models and durable evidence

I benchmark smaller local models on repeatable engineering tasks and build durable execution records for evaluation and future specialization experiments.

The problem

General-purpose frontier models can be an expensive default for repeatable tasks. Choosing a smaller model requires evidence about task quality, context size, latency, quantization, runtime, and the cost of correcting failures.

Scope

On a single AMD Strix Halo machine with 128 GB unified memory, I run local inference and compare context size, quantization, runtimes, quality, and routing against hosted options. Structured execution evidence is accumulating for evaluation; LoRA and specialist training and deployment are future investigations.

My contribution

I am shaping evaluation and data boundaries that make model comparisons actionable: validated execution becomes durable evidence I can use to inform routing and future specialist-model experiments.

Architecture and technical decisions

  1. Route by task difficulty, privacy, cost, latency, and demonstrated quality instead of sending every task to the largest available model.

  2. Benchmark small and mid-sized models against the same task evidence and validation criteria used for frontier-model comparisons.

  3. Keep execution history structured and searchable so validated traces can support failure analysis, evaluation, and future training datasets.

  4. Treat LoRA and specialist-model training as a future path that depends on the quality and coverage of accumulated evidence.

Constraints and qualifications

  • Model quality must be measured against the same tasks and validation criteria before routing changes.
  • Context size, quantization, runtime, latency, privacy, and correction cost affect the useful choice.
  • Local benchmarks run on one AMD Strix Halo machine with 128 GB unified memory; structured execution evidence is accumulating now.
  • Future investigations include LoRA training and CPU-friendly specialist deployment for classification, bounded decisions, completion checks, and narrow reviewer or evaluator tasks.

Guiding principles

  • Use the least expensive intelligence that reliably solves a task.
  • Retain validated execution history as an evaluation and dataset asset, not disposable logs.
  • Route only when evidence shows a model meets the task quality and operating constraints.

System views

Choose intelligence by the work and its constraints.

I benchmark local and hosted models against task quality, context, latency, and cost. One AMD Strix Halo machine with 128 GB unified memory handles my primary local-AI workload. Specialist models remain a future path under evaluation.

Models — conceptual system topologyAn architectural illustration, not live telemetry or model availability. Prefer deterministic software or a small capable model; use frontier reasoning when the task warrants it, based on quality and constraints. Execution mode: once. Components: Engineering task (Difficulty · failure cost); Assess constraints (Privacy · context · latency); Compare evidence (Quality · cost · reliability); Deterministic software (Known structure · rules); Small local model (Private · fast inference); Specialist model (Future · after evaluation); Frontier model (Complex reasoning); Human escalation (High risk · uncertainty); Validated outcome (Route from evidence). Ordered stages: Engineering task; Assess constraints; Compare evidence; Deterministic software and Small local model and Frontier model; Validated outcome. Connections: Engineering task to Assess constraints; Assess constraints to Compare evidence; Compare evidence to Deterministic software; Compare evidence to Small local model; Compare evidence to Specialist model (conditional); Compare evidence to Frontier model; Compare evidence to Human escalation (when risk requires) (conditional); Deterministic software to Validated outcome; Small local model to Validated outcome; Specialist model to Validated outcome (conditional); Frontier model to Validated outcome; Human escalation to Validated outcome (conditional). Optional conditional paths are shown but do not run in the illustrated sequence.Engineering taskDifficulty · failure costAssess constraintsPrivacy · context ·latencyCompare evidenceQuality · cost ·reliabilityDeterministicsoftwareKnown structure · rulesSmall local modelPrivate · fast inferenceSpecialist modelFuture · after evaluationFrontier modelComplex reasoningHuman escalationHigh risk · uncertaintyValidated outcomeRoute from evidence
  • idle
  • incoming
  • active
  • outgoing
  • settled

Conceptual sequence · not live telemetry.

Execution sequence

  1. 01
    Engineering taskEngineering task → Assess constraints
  2. 02
    Assess constraintsAssess constraints → Compare evidence
  3. 03
    Compare evidenceCompare evidence → Deterministic software · Compare evidence → Small local model · Compare evidence → Frontier model
  4. 04
    Deterministic software + Small local model + Frontier modelDeterministic software → Validated outcome · Small local model → Validated outcome · Frontier model → Validated outcome
  5. 05
    Validated outcome

Conditional paths

  • Compare evidence → Specialist model · Future · after evaluation
  • Compare evidence → Human escalation · High risk · uncertainty · when risk requires
  • Specialist model → Validated outcome
  • Human escalation → Validated outcome
Conceptual execution · explicit dependencies · not live telemetry

Every run can improve the next system decision.

I am building evidence capture into my own agent workflows to record task context, tool activity, candidate results, validation, independent review, and outcomes from agent work.

Evidence and learning — conceptual system topologyAn architectural illustration, not live telemetry. Structured execution history supports evaluation now and offers a future path to specialist-model datasets. Execution mode: continuous. Components: Task definition (Goal · constraints); Context used (Sources · provenance); Tool activity (Calls · outputs); Candidate result (Proposed change); Validation (Tests · deterministic checks); Independent review (Findings · confidence); Human decision (Escalation · acceptance); Final outcome (Verified · stopped · revised); Durable evidence (Structured execution history); Workflow evaluation (Reliability · failure analysis); Model routing (Quality · cost · latency); Future specialization (Future training · optional). Ordered stages: Task definition; Context used; Tool activity; Candidate result; Validation; Independent review; Final outcome; Durable evidence. Feedback loop: Durable evidence; Workflow evaluation; Model routing; the next continuous pass explicitly revisits Context used; Tool activity; Candidate result; Validation; Independent review; Final outcome; Durable evidence. Connections: Task definition to Context used; Context used to Tool activity; Tool activity to Candidate result; Candidate result to Validation; Validation to Independent review; Independent review to Human decision (when review escalates) (conditional); Human decision to Final outcome (conditional); Independent review to Final outcome; Final outcome to Durable evidence; Durable evidence to Workflow evaluation (feedback); Workflow evaluation to Model routing (feedback); Durable evidence to Future specialization (conditional); Model routing to Context used (feedback). Optional conditional paths are shown but do not run in the illustrated sequence.Task definitionGoal · constraintsContext usedSources · provenanceTool activityCalls · outputsCandidate resultProposed changeValidationTests · deterministicchecksIndependent reviewFindings · confidenceHuman decisionEscalation · acceptanceFinal outcomeVerified · stopped ·revisedDurable evidenceStructured executionhistoryWorkflow evaluationReliability · failureanalysisModel routingQuality · cost · latencyFuture specializationFuture training · optional
  • idle
  • incoming
  • active
  • outgoing
  • settled

Conceptual sequence · not live telemetry.

Execution sequence

  1. 01
    Task definitionTask definition → Context used
  2. 02
    Context usedContext used → Tool activity
  3. 03
    Tool activityTool activity → Candidate result
  4. 04
    Candidate resultCandidate result → Validation
  5. 05
    ValidationValidation → Independent review
  6. 06
    Independent reviewIndependent review → Final outcome
  7. 07
    Final outcomeFinal outcome → Durable evidence
  8. 08
    Durable evidence

Ongoing learning loop

  1. 01
    Durable evidenceDurable evidence → Workflow evaluation
  2. 02
    Workflow evaluationWorkflow evaluation → Model routing
  3. 03
    Model routingModel routing → Context used

The next pass explicitly revisits Context used → Tool activity → Candidate result → Validation → Independent review → Final outcome → Durable evidence for another pass.

Conditional paths

  • Independent review → Human decision · Escalation · acceptance · when review escalates
  • Human decision → Final outcome
  • Durable evidence → Future specialization · Future training · optional
Conceptual execution · explicit dependencies · not live telemetry

Intended effect

Current work accumulates validated execution evidence and benchmarks smaller models. Future investigations include CPU-friendly specialists for classification, bounded decisions, completion checking, and narrow reviewer or evaluator tasks, with LoRA training and deployment informed by those records.

Technologies and components

  • PostgreSQL
  • Vector retrieval
  • Local inference
  • Quantization
  • LoRA (future study)