Selected system / personal and professional work
Experiment in progressSpecialized local models and durable evidence
I benchmark smaller local models on repeatable engineering tasks and build durable execution records for evaluation and future specialization experiments.
The problem
General-purpose frontier models can be an expensive default for repeatable tasks. Choosing a smaller model requires evidence about task quality, context size, latency, quantization, runtime, and the cost of correcting failures.
Scope
On a single AMD Strix Halo machine with 128 GB unified memory, I run local inference and compare context size, quantization, runtimes, quality, and routing against hosted options. Structured execution evidence is accumulating for evaluation; LoRA and specialist training and deployment are future investigations.
My role
My contribution
I am shaping evaluation and data boundaries that make model comparisons actionable: validated execution becomes durable evidence I can use to inform routing and future specialist-model experiments.
System design
Architecture and technical decisions
Route by task difficulty, privacy, cost, latency, and demonstrated quality instead of sending every task to the largest available model.
Benchmark small and mid-sized models against the same task evidence and validation criteria used for frontier-model comparisons.
Keep execution history structured and searchable so validated traces can support failure analysis, evaluation, and future training datasets.
Treat LoRA and specialist-model training as a future path that depends on the quality and coverage of accumulated evidence.
System boundaries
Constraints and qualifications
- Model quality must be measured against the same tasks and validation criteria before routing changes.
- Context size, quantization, runtime, latency, privacy, and correction cost affect the useful choice.
- Local benchmarks run on one AMD Strix Halo machine with 128 GB unified memory; structured execution evidence is accumulating now.
- Future investigations include LoRA training and CPU-friendly specialist deployment for classification, bounded decisions, completion checks, and narrow reviewer or evaluator tasks.
Design approach
Guiding principles
- Use the least expensive intelligence that reliably solves a task.
- Retain validated execution history as an evaluation and dataset asset, not disposable logs.
- Route only when evidence shows a model meets the task quality and operating constraints.
Conceptual architecture
System views
Choose intelligence by the work and its constraints.
I benchmark local and hosted models against task quality, context, latency, and cost. One AMD Strix Halo machine with 128 GB unified memory handles my primary local-AI workload. Specialist models remain a future path under evaluation.
- idle
- incoming
- active
- outgoing
- settled
Conceptual sequence · not live telemetry.
Execution sequence
- 01Engineering taskEngineering task → Assess constraints
- 02Assess constraintsAssess constraints → Compare evidence
- 03Compare evidenceCompare evidence → Deterministic software · Compare evidence → Small local model · Compare evidence → Frontier model
- 04Deterministic software + Small local model + Frontier modelDeterministic software → Validated outcome · Small local model → Validated outcome · Frontier model → Validated outcome
- 05Validated outcome
Conditional paths
- Compare evidence → Specialist model · Future · after evaluation
- Compare evidence → Human escalation · High risk · uncertainty · when risk requires
- Specialist model → Validated outcome
- Human escalation → Validated outcome
Every run can improve the next system decision.
I am building evidence capture into my own agent workflows to record task context, tool activity, candidate results, validation, independent review, and outcomes from agent work.
- idle
- incoming
- active
- outgoing
- settled
Conceptual sequence · not live telemetry.
Execution sequence
- 01Task definitionTask definition → Context used
- 02Context usedContext used → Tool activity
- 03Tool activityTool activity → Candidate result
- 04Candidate resultCandidate result → Validation
- 05ValidationValidation → Independent review
- 06Independent reviewIndependent review → Final outcome
- 07Final outcomeFinal outcome → Durable evidence
- 08Durable evidence
Ongoing learning loop
- 01Durable evidenceDurable evidence → Workflow evaluation
- 02Workflow evaluationWorkflow evaluation → Model routing
- 03Model routingModel routing → Context used
The next pass explicitly revisits Context used → Tool activity → Candidate result → Validation → Independent review → Final outcome → Durable evidence for another pass.
Conditional paths
- Independent review → Human decision · Escalation · acceptance · when review escalates
- Human decision → Final outcome
- Durable evidence → Future specialization · Future training · optional
Intended effect
Intended effect
Current work accumulates validated execution evidence and benchmarks smaller models. Future investigations include CPU-friendly specialists for classification, bounded decisions, completion checking, and narrow reviewer or evaluator tasks, with LoRA training and deployment informed by those records.
Technologies and components
- PostgreSQL
- Vector retrieval
- Local inference
- Quantization
- LoRA (future study)