Michael Puodziukas, AI evaluation engineer, Remote (US)
Replay a synthetic COBOL payroll defect yourself: deterministic, public, checkable.
I build tests that catch AI model failures before release, and publish them so anyone can run them.
Open to mid-level and senior IC roles, W2, remote (US).

A COBOL PICTURE-clause defect, reproduced on a public synthetic probe: 47,312 records, $40,812.81 loss, 0 divergence.
I build and red-team AI systems in the same pass: private inference, adversarial gates, deterministic replay. The failure class surfaces in the build, not production. Remote, async-first.
The headline replays from a 47,312-record deterministic corpus. Clone and run: cobol-pic-probe · llm-adversarial-gate · exo
$ git clone https://github.com/mpuodziukas-labs/cobol-pic-probe && cd cobol-pic-probe && pytestCOBOL PIC 9(7)V99 fixed-point truncation, synthetic, fixed-seed, reproducible.
Offline adversarial-validation gate, mapped to the OWASP LLM Top 10.
An OWASP checklist lists known risks; llm-adversarial-gate runs labeled attacks against them.
140/140 caught, 0 bypasses on a synthetic corpus a checklist never runs.
An OWASP checklist confirms known-issue coverage. An adversarial run finds what the checklist can't. llm-adversarial-gate is the public anchor: fixed corpus, deterministic, clone-and-run.
Verify, don't trust
Three systems: two public and clone-and-run, one a public fork hardened in the home lab. More on /work.
COBOL forensics, adversarial-AI defense, self-hosted inference. Clone the public ones and run them yourself.
COBOL · LLM Security · Inference
COBOL Payroll Forensics
COBOL Forensic Audit (methodology reference)
A synthetic COBOL payroll corpus reproduces a silent COMP-3 fixed-point truncation, a class that standard audits and monitoring can miss. Public, deterministic, clone and run.
Failure class: PICTURE-clause fixed-point truncation, compounding below the totals-reconciliation threshold.
Full teardown →
Silent COMP-3 arithmetic overflow on a synthetic, fixed-seed payroll corpus. COMP-3 PIC clause mismatch: ON SIZE ERROR suppressed, silent truncation, wrong multiplier on overtime. Detected via adversarial boundary scanner. Repeated runs return identical results (test_repeated_runs_are_identical). Zero divergence across 47,312 synthetic records, $40,812.81 synthetic loss. Failure class: PICTURE-clause fixed-point truncation compounding below totals-reconciliation thresholds. Root cause documented, replay-verified. Every claim has a replay command: pytest cobol-pic-probe/ -v
$ python3 probe.py --records 47312LLM Adversarial Validation Council
Adversarial AI validation · own synthetic test corpus, n=220 (clone-and-run)
A rule-based gate checks every LLM output before it ships, with a human-in-the-loop review step. Result on its own synthetic test corpus (n=220, written for the repo): 140/140 adversarial attempts caught, 0/80 benign over-blocked. Clone and run.
Baseline: no adversarial gate on LLM output means every wrong answer ships undetected.
Full teardown →
An LLM pipeline with no adversarial gate in place ships wrong answers with no check and no replay trail. A rule-based gate, not a policy document: a working deterministic check, HITL review step, CI green. Result: 140/140 adversarial attempts caught on its own synthetic test corpus (n=220, written for the repo), 0/80 benign over-blocked, mapped to the OWASP LLM Top 10. Every attack attempt logged with vector classification and a replay artifact, replayable not asserted.
$ python3 -m pytest tests/ -vSelf-Hosted Inference (Home Lab)
Open Source + Self-built · my public fork of exo (mpuodziukas-labs/exo), ops-hardening branch
My public fork of exo (open-source distributed inference), with an ops-hardening branch, hardened in the home lab. Self-hosted inference first; cloud is the stated fallback.
Managed-API inference means a single-vendor dependency and every query leaving the building.
Full teardown →
Managed-API inference means a single-vendor dependency and per-query data egress. I built a self-hosted alternative in the home lab. My public fork of exo (open-source distributed inference), with an ops-hardening branch. Infrastructure-as-code: Terraform on ARM. mTLS everywhere. Every routing decision logged with model and confidence.
Methodology: Every public metric clones and runs. The replay corpus (0 divergence) is public and CI-green: github.com/mpuodziukas-labs/cobol-pic-probe. Replay determinism: repeated runs return identical results (test_repeated_runs_are_identical).
Where this breaks
Every claim above names its own edge.
The inference cluster is peer-loss survivable, not fault-free.
Two nodes shard the model running in the home lab. If the primary node drops the peer, it serves solo; if the primary node itself is lost, the peer cannot hold the model and service stops. Asymmetric fallback, stated plainly.
The home-lab model is a quantized class, not a frontier model.
What I claim is reproducible self-hosted inference on hardware I own, at a known cost. I do not claim the model matches a frontier model.
Work with me
The systems above are public code you can clone and run.
Three systems ship in public, reproducible from public code. Open to mid-level and senior IC roles · full-time W2 · Remote (US-based).
Work with me →