Project “Angelina”

A zero-cost, self-hosted, hybrid-cloud AI system — engineered for reliability under hard constraints.

$0 sustained cost No production regressions
What it is

A resilient two-node AI agent that costs nothing to run

A self-hosted AI agent (Python + FastAPI, a hosted LLM on its free tier, a vector database for retrieval, and SQLite for state) that performs a scheduled daily task, keeps a retrievable knowledge base in sync, and serves a phone-friendly chat interface over HTTPS. The interesting part isn't the model call — it's the operations engineering around it: a hybrid local-primary / cloud-passive deployment with heartbeat-based failover, authenticated bi-directional sync, and fully idempotent infrastructure-as-code, all delivered inside always-free tiers without ever breaking the running service.

Architecture

Local does the work. Cloud stays quiet — until it can't.

01 · PRIMARY

Local VM

Home-lab VM runs the app under systemd and fires the daily task on a cron schedule. All real work happens here.

02 · PASSIVE

Cloud backup

An always-free cloud VM sits idle, bound to localhost behind HTTPS, ready to take over only when needed.

03 · COORDINATION

Heartbeat failover

Local stamps a presence heartbeat; cloud checks it and fails over only if local goes silent past a 25-hour window.

04 · DATA

HTTPS sync

A token-gated endpoint exchanges knowledge vectors and history daily, merged last-write-wins, re-embedded on import.

Show the full topology (ASCII)
                      shared sheet (heartbeat relay + logging)
                      ▲         ▲
      stamp heartbeat │         │ read heartbeat at 23:00 check
                      │         │
   ┌──────────────────┴──┐   ┌──┴───────────────────┐
   │  LOCAL  (PRIMARY)    │   │  CLOUD (PASSIVE)      │
   │  home-lab VM         │   │  always-free e2-micro │
   │  • FastAPI (LAN)     │   │  • app bound to       │
   │  • cron daily task   │   │    127.0.0.1          │
   │  • heartbeat (local  │   │  • Caddy -> HTTPS      │
   │    role only)        │   │  • dynamic DNS        │
   └──────────┬───────────┘   │  • 22/80/443 only,    │
              │               │    SSH key-only       │
              │  token-gated  └──────────┬────────────┘
              │  HTTPS sync (last-write-  │
              └──  wins, re-embed) ───────┘

   remote user ──HTTPS──▶ Caddy ──▶ cloud app ──▶ chat UI + /audit
Engineering highlights

What makes it reliable

Failover

At-most-once failover

Heartbeat + 25-hour presence window guarantees exactly one daily run across two nodes — never a duplicate.

Sync

Bi-directional HTTPS sync

Token-gated daily exchange of knowledge vectors and history, last-write-wins, re-embedded on import.

IaC

Idempotent Ansible

A single playbook provisions the whole host — timezone, swap, firewall, runtime, app, TLS — re-runnable any time.

Ops

Audit & self-healing

Read-only security audit scoring, gated service self-healing with a preview mode, and proactive TLS-expiry checks.

Access

Secure HTTPS remote access

Auto certificates via reverse proxy, constant-time token auth, SSH key-only, root disabled, minimal open ports.

Cost

Zero cost by design

Every component lives inside an always-free tier, with a $1 budget tripwire that has never fired.

War stories

Nine production bugs, nine lessons

Reliability was earned the hard way. Each bug taught a generalizable lesson — full write-up in the post-mortem.

1. Credential clobbering on deploy

Config is not code — keep secrets out of source.

2. Cron ran without the environment

cron inherits nothing; load your environment explicitly.

3. Python 3.11 vs 3.12 syntax drift

“Works on my machine” has a version number attached.

4. SQLite upsert vs partial index

Upsert arbiters have precise, engine-specific rules.

5. Namespace-bounce duplication

In sync, identity must be stable and origin-aware.

6. Synchronous re-embed timeout

Keep slow, unbounded work off the request path.

7. Failover double-stamp

A presence signal must be written only by its owner.

8. Transient 503 missed a push

Retry with capped backoff for rare, high-stakes jobs.

9. UTC vs local timezone cron

Set the timezone explicitly; schedulers trust the clock blindly.

Read the full post-mortem →

Results

Measured outcomes

$0
sustained cost; tripwire never fired
Multi-day
failover verified in production
At-most-once
no duplicate daily output
v1 → v3.4.1
disciplined, secret-scanned releases