Skip to content

AI Lab / my personal workshop

Systems I build, operate, and test.

The lab exists to compress the engineering feedback loop. I build and operate software and AI systems daily in my own environment, moving from an idea to real use and revision—sometimes in hours or days. That experience informs my enterprise architecture work, where security, governance, reliability, and team ownership serve a different scale.

AI Lab / environment

The environment where my AI works.

Much of this environment is in daily use. I build the surrounding infrastructure as well as the AI workflows; nearly all AI-assisted development runs in a dedicated server-side sandbox.

Isolated execution

The sandbox isolates agent work with controlled access to repositories, tooling, and credentials. It routes hosted and local models into bounded, repeatable tasks so I can replace tooling without destabilizing my workstation.

Models and compute

One AMD Strix Halo machine with 128 GB unified memory supports local inference and workers. I benchmark context size, quantization, runtime, task quality, and routing against hosted options. The practical question is which useful work smaller local systems can handle reliably.

The supporting stack

I run VMs, databases, analytics, monitoring, storage and backups alongside applications and services. Deployment and recovery belong to the same operating stack.

Checks and evidence

Orchestration, validation, and review give bounded agent work a path to inspectable results. I am building durable evidence capture to compare outcomes and revise the workflow. Structured execution evidence is accumulating for evaluation and future specialist experiments.

AI Lab / operating loop

From idea to a better next version.

I do this because architecture, infrastructure, AI, developer tooling, and product engineering are the parts of engineering I most enjoy—and this environment lets me work across all of them.

Idea → architecture → implementation → deployment → real use → evidence → opinion → revision. Independent decisions, hands-on operations, and rapid replacement make that loop useful in my own environment.

I start with a small deterministic tool where it is enough, use a capable local model where it holds up under validation, and call on frontier reasoning when the problem warrants it. The cards below show which work is built and used, being benchmarked or tested, being built, or still a future investigation.

AI Lab / systems and experiments

Working systems and experiments.

Built work, benchmarks, active tests, and future investigations each have their own stage. Results from real use inform the next iteration and my professional judgment.

  • TOPIC 01
    Benchmarking

    Local model performance and routing

    I benchmark local inference on my AMD Strix Halo machine with 128 GB unified memory, comparing context, quantization, runtimes, latency, and task quality before changing routing.

  • TOPIC 02
    Built and used

    Deterministic code exploration

    I built and use indexes, syntax-aware analysis, repository graphs, and targeted retrieval to find relevant code before model reasoning.

  • TOPIC 03
    Testing

    Context-size experiments

    I am testing how context selection and size affect task quality, completeness, latency, and cost.

  • TOPIC 04
    Building now

    Agent evaluation

    I am building repeatable tasks, validation checks, and structured evidence to compare agent workflows and model choices.

  • TOPIC 05
    Testing

    Independent model review

    I am testing how an independent reviewer can inspect candidate changes against task context, constraints, and validation results.

  • TOPIC 06
    Building now

    DAG-based orchestration

    I am building orchestration that makes dependencies, parallel work, retries, validation gates, and human approval visible across deterministic and AI-driven steps.

  • TOPIC 07
    Building now

    Evidence-driven AI workflows

    I am building durable capture of task context, tool activity, candidate output, validation, independent review, and outcomes as inspectable evidence.

  • TOPIC 08
    Testing

    RAG and vector retrieval

    I am testing retrieval as a bounded context source with provenance, clear service ownership, and evidence about relevance.

  • TOPIC 09
    Future investigation

    Specialist models and LoRA

    Future investigation: use accumulated validated execution evidence to evaluate LoRA and potentially train CPU-friendly specialists for classification, bounded decisions, completion checks, and narrow reviewer or evaluator work.