NEXUS — Autonomous Homelab SRE
An evidence-first autonomous SRE agent built around four hard rules — evidence over vibes, allowlisted typed tools, approval enforced outside the model, and explainable reasoning. Investigates incidents the way a careful operator would, verified live against PostgreSQL and Redis.
What was happening.
An AI SRE for real infrastructure that has to show its work: it observes the live system, collects evidence with typed tools, forms hypotheses, tests them, and only then explains a root cause — and it refuses to sound confident when it cannot prove something.
What needed to change.
- Most “AI ops” demos ask an LLM to guess and let confidence stand in for correctness.
- Without structured tools and permissions, agents either do nothing useful or have an alarming amount of power — arbitrary bash, rm -rf — with no audit trail.
- Explainability gets sacrificed: conclusions arrive without observable evidence or a reasoning path you can verify.
How it was done.
Confidence is derived from observations, contradictions and required-evidence coverage — never invented by the model. When NEXUS cannot verify something, the honest output is “I could not verify disk usage because the host did not respond” — not “disk usage is normal.”
No arbitrary shell. Every capability is a typed tool with a permission class (READ_ONLY, LOW_RISK, REQUIRES_APPROVAL, FORBIDDEN). The model cannot bypass the permission layer because authorization is enforced outside the LLM.
REQUIRES_APPROVAL actions cannot run until a human says yes. Default mode is READ_ONLY — no destructive operations until the permission system is explicitly enabled.
The console shows reasoning summaries and the evidence collected at every step — never hidden chain-of-thought.
classify → context → evidence → hypothesis → test → root cause → remediation → approval → execute → verify → close. When evidence is insufficient the reasoning loop iterates instead of inventing a root cause just to finish.
FastAPI application factory with a live /api/v1/health database check, PostgreSQL 16 persistence (SQLAlchemy 2.0 async, Alembic migrations), Redis, structured logging with structlog, a sandbox Docker Compose stack, and CI running ruff, strict mypy and pytest against real PostgreSQL and Redis — with graceful degradation when a dependency is down.
The numbers that came out.
No destructive ops until the permission system is enabled.
READ_ONLY, LOW_RISK, REQUIRES_APPROVAL, FORBIDDEN.
Verified against live PostgreSQL 16 + Redis 7.
Hypotheses tested against the live system before claiming anything.
Where it landed.
- An agent architecture where sounding confident is never confused with being correct — the core fix for the AI-ops hype problem.
- Safety enforced structurally: READ_ONLY by default, four permission classes, human approval outside the model, and a complete audit trail.
- Phase 1 foundation verified against live PostgreSQL and Redis; subsequent phases — discovery, typed tools, LangGraph agent, evaluation framework — documented and scheduled in verifiable milestones.
- An honest feature ledger: everything is marked REAL, SIMULATED, MOCK or NOT IMPLEMENTED — nothing is claimed before it is tested.
More case studies.
Full-Stack School ERP
Designed and shipped a full-stack ERP end-to-end — from requirements through deployment — consolidating student records, attendance, finance and communication, with notification, messaging and authentication capabilities serving the whole institution.
Read case studyAtlas — Distributed Infrastructure Engine
Atlas discovers and models infrastructure as a graph, schedules workloads, replicates state through a from-scratch Raft layer, injects controlled failures, correlates them into incidents and produces advisory AI analysis — running entirely over an in-memory event bus on a single machine, with real-time visualization and algorithmic transparency.
Read case studyCS2RGB — Counter-Strike × OpenRGB Lighting
A Python service that taps CS2's Game State Integration API and drives OpenRGB directly — lighting reacts to health, flash/smoke/burn status, game phase and round events (bomb, kills, round wins) with secure secret-key authentication and full event logging.
Read case studyYour situation probably looks different — but the method doesn't.
Understand the work, decide the approach, deliver and measure. A free conversation establishes whether it applies to your problem.
or email joseph.gitau.c@gmail.com