Skip to main content
Hiring for:AI Evaluation EngineerAI Red-Team / Safety Evaluation EngineerAI Governance & GRC Engineer

Michael Puodziukas, AI evaluation engineer, Remote (US)

Replay a synthetic COBOL payroll defect yourself: deterministic, public, checkable.

I build tests that catch AI model failures before release, and publish them so anyone can run them.

Work with meRésumé (PDF)

Open to mid-level and senior IC roles, W2, remote (US).

Topographic data-mesh backdrop in navy and royal blue

A COBOL PICTURE-clause defect, reproduced on a public synthetic probe: 47,312 records, $40,812.81 loss, 0 divergence.

I build and red-team AI systems in the same pass: private inference, adversarial gates, deterministic replay. The failure class surfaces in the build, not production. Remote, async-first.

The headline replays from a 47,312-record deterministic corpus. Clone and run: cobol-pic-probe · llm-adversarial-gate · exo

An OWASP checklist lists known risks; llm-adversarial-gate runs labeled attacks against them.

140/140 caught, 0 bypasses on a synthetic corpus a checklist never runs.

An OWASP checklist confirms known-issue coverage. An adversarial run finds what the checklist can't. llm-adversarial-gate is the public anchor: fixed corpus, deterministic, clone-and-run.

Verify, don't trust

Three systems: two public and clone-and-run, one a public fork hardened in the home lab. More on /work.

COBOL forensics, adversarial-AI defense, self-hosted inference. Clone the public ones and run them yourself.

COBOL · LLM Security · Inference

47,312 synthetic replays, fixed seed
synthetic payroll corpus

COBOL Payroll Forensics

COBOL Forensic Audit (methodology reference)

A synthetic COBOL payroll corpus reproduces a silent COMP-3 fixed-point truncation, a class that standard audits and monitoring can miss. Public, deterministic, clone and run.

COBOLPythonReplay DeterminismAdversarial Boundary Analysis

Failure class: PICTURE-clause fixed-point truncation, compounding below the totals-reconciliation threshold.

Full teardown →

Silent COMP-3 arithmetic overflow on a synthetic, fixed-seed payroll corpus. COMP-3 PIC clause mismatch: ON SIZE ERROR suppressed, silent truncation, wrong multiplier on overtime. Detected via adversarial boundary scanner. Repeated runs return identical results (test_repeated_runs_are_identical). Zero divergence across 47,312 synthetic records, $40,812.81 synthetic loss. Failure class: PICTURE-clause fixed-point truncation compounding below totals-reconciliation thresholds. Root cause documented, replay-verified. Every claim has a replay command: pytest cobol-pic-probe/ -v

PIC boundary scanner → COMP-3 overflow trace → deterministic replay harness, fixed random seed, repeated runs return identical output
03:14→INPUT TRACE— PIC V9 truncation: full corpus silently truncated. ON SIZE ERROR suppressed.
04:00→PROOF— PICTURE-clause truncation confirmed. Deterministic replay: repeated runs return identical results.
DEPLOY→VERIFIED— Zero variance. test_repeated_runs_are_identical passes on every run.
CI: tests passing$ python3 probe.py --records 47312
Clone & run on GitHub →
0 successful bypasses · own test corpus
OWASP LLM Top 10 mapped

LLM Adversarial Validation Council

Adversarial AI validation · own synthetic test corpus, n=220 (clone-and-run)

A rule-based gate checks every LLM output before it ships, with a human-in-the-loop review step. Result on its own synthetic test corpus (n=220, written for the repo): 140/140 adversarial attempts caught, 0/80 benign over-blocked. Clone and run.

PythonHITLQdrantFastEmbed

Baseline: no adversarial gate on LLM output means every wrong answer ships undetected.

Full teardown →

An LLM pipeline with no adversarial gate in place ships wrong answers with no check and no replay trail. A rule-based gate, not a policy document: a working deterministic check, HITL review step, CI green. Result: 140/140 adversarial attempts caught on its own synthetic test corpus (n=220, written for the repo), 0/80 benign over-blocked, mapped to the OWASP LLM Top 10. Every attack attempt logged with vector classification and a replay artifact, replayable not asserted.

Input → rule-based adversarial gate → HumanGate → output
BASELINE→NO GATE— No adversarial gate on LLM output. Every wrong answer ships unflagged.
BUILT→GATE WIRED— Rule-based adversarial gate wired, HITL review step in place.
RESULT→CI GREEN— 140/140 adversarial caught on its own synthetic test corpus (n=220, written for the repo). 0/80 benign over-blocked. 114 offline pytest assertions.
CI: tests passing$ python3 -m pytest tests/ -v
Clone & run on GitHub →
Public fork of exo, ops-hardening branch
ops-hardening branch

Self-Hosted Inference (Home Lab)

Open Source + Self-built · my public fork of exo (mpuodziukas-labs/exo), ops-hardening branch

My public fork of exo (open-source distributed inference), with an ops-hardening branch, hardened in the home lab. Self-hosted inference first; cloud is the stated fallback.

PythonMLXTerraformmTLS

Managed-API inference means a single-vendor dependency and every query leaving the building.

Full teardown →

Managed-API inference means a single-vendor dependency and per-query data egress. I built a self-hosted alternative in the home lab. My public fork of exo (open-source distributed inference), with an ops-hardening branch. Infrastructure-as-code: Terraform on ARM. mTLS everywhere. Every routing decision logged with model and confidence.

Multi-tier routing: self-hosted before cloud, cloud as fallback.
AUDIT→DEPENDENCY— Managed-API dependency audited. Single-vendor lock, per-query data egress.
REBUILD→SELF-HOSTED— exo + MLX self-hosted in the home lab.
STEADY→MAINTAINED— Ops-hardening branch maintained, KV-cache tuned on every path.
View Repository →

The full proof record · puodziukas.dev/proof

Methodology: Every public metric clones and runs. The replay corpus (0 divergence) is public and CI-green: github.com/mpuodziukas-labs/cobol-pic-probe. Replay determinism: repeated runs return identical results (test_repeated_runs_are_identical).

Where this breaks

Every claim above names its own edge.

  • The inference cluster is peer-loss survivable, not fault-free.

    Two nodes shard the model running in the home lab. If the primary node drops the peer, it serves solo; if the primary node itself is lost, the peer cannot hold the model and service stops. Asymmetric fallback, stated plainly.

  • The home-lab model is a quantized class, not a frontier model.

    What I claim is reproducible self-hosted inference on hardware I own, at a known cost. I do not claim the model matches a frontier model.

Work with me

The systems above are public code you can clone and run.

Three systems ship in public, reproducible from public code. Open to mid-level and senior IC roles · full-time W2 · Remote (US-based).

Work with me →