AHMAD:// Available for work — open to remote roles and contracts
fetched yesterday

Muhammad Ahmad

I build agentic AI systems with enforceable trust boundaries.

Runtimes, evals and audits — spend ceilings agents can't overrun, policies that fail closed, logs you can verify offline. Everything on this site links to its proof; CI re-checks those links nightly and admits it when one rots.

HOW I BUILD AI-assisted development, human-verified: I write the contracts, review every line, tests decide. the manifesto →

claims → evidence · sample

full ledger: run `evidence` — every row opens something you can check yourself
claimstatusevidence
The agent loop cannot overrun a token or dollar ceiling — enforced in code and covered by tests TESTED test/cost.test.ts
Publicly installable from npm as @rrrtx/xr PUBLISHED npmjs.com/package/@rrrtx/xr
Forty-nine-finding forensic self-audit run before release; every finding closed with an evidence note AUDITED docs/audit/FINAL_AUDIT.md

> ls ~/featured

01Live · active

xr

A local-first AI agent runtime that cannot overspend its budget or hide its actions.

721commits 65merged PRs CI workflows v3.1.5on npm

TypeScript · Bun · SQLite · MCP

02Active

xr-foundation-model

A tiny GPT grown from zero — tokenizer, training loop, math proofs — and then audited by me, in public.

  • 49-finding forensic self-audit, closed 49→0
  • val perplexity 301.55 — measured, then admitted it's small
  • checkpoint committed to the repo as evidence (76.8 MB)
  • 33 test files incl. ground-truth math proofs

Python · PyTorch · from-scratch BPE · pytest

03

commit-canvas

Live

Turn any git repository into an animated, shareable story page. Zero cost, zero auth, one command. Starred by three early users, its own Actions pipeline renders this repo's story page on every push — dogfooding included.

Python · HTML/JS · GitHub Actions · 0 network requests

> view all work

real output · every line linked to its source

No fake logs, no invented telemetry. This is a replay of what my own release audit and CI actually recorded — failures included, because that's the point of keeping them visible.

ahmad@xr:~% — session replay · all output real · typed once per visit
$ bun test
238 files · 2924 pass · 13 skip · 0 fail · 12,978 expects · 57.6 s
$ curl -fsSL https://raw.githubusercontent.com/ahmadrrrtx/xr/main/install.sh | bash
binary installed · @rrrtx/xr v3.1.5 · MIT · golden path 0.76 s e2e
$ gh run view 31702859699 # cross-platform CI — as my own audit documents it
✘ linux / macos / windows failed at step 7 “Full unit suite”
real, environment-specific failures (100 s / 93 s / 326 s) — not timeouts, not hidden
$ cat docs/audits/XR_FINAL_RELEASE_AUDIT.md | grep -i verdict
✔ “No evidence of fabricated test success or vacuous coverage was found.”
watch · repo + verifier snapshot baked at build — refreshes on every push
git log -1 xr@ea21bd4 "Merge pull request #75 from ahmadrrrtx/phase2/decisi…" · yesterday
fetched yesterday
roleAgentic AI Engineer
focusruntimes · evals · audits
status● open to work
locationPakistan · UTC+5
latest@rrrtx/xr v3.1.5
links checkednightly →
host Muhammad-Ahmad · Chiniot, PK
shells zsh, bash (both real)
editor neovim + AI agents — disclosed
uptime 2 focused hours/day, ~7d/wk
focus Agents / LLMs / Systems
claims 0 unverified on this site
  • 2026-09First external open-source PRs — evaluating promptfoo's eval API while writing my own golden-dataset harness.
  • 2026-09xr: decision-boundary hardening phases merging on branches; the audit-chain spec under review.
  • 2026-09AegisRoute rebuild: same idea, xr-grade CI discipline. One-evening repos are the pattern I'm breaking.

the honest version of this list →

> cat ~/writing

email → ahmadrrrtx@gmail.com github → github.com/ahmadrrrtx · client work → rrrtx-systems.com writing → ahmadrrrtx.substack.com ↗ — newsletter, occasional engineering notes, unsubscribe anytime also on · substack ↗· medium ↗· hashnode ↗· dev.to ↗· indie hackers ↗· daily.dev ↗
open a conversation
ahmad@xr:~% ask — grounded in this site · no LLM, no invented answers

try: