EC2 source-code discovery and reliability triage (#2)

Second discovery pass over the Scrivas EC2 estate (`716468089330`, `us-east-2`, 7 instances). Prompted by two client signals: they do not hold the source code for their own platform (built under contract by the incumbent vendor), and they are reporting production reliability issues.

All access was **read-only** — no writes, restarts or config changes on any instance.

## Source code is recoverable

**All 10 application repositories exist as complete git checkouts on Scrivas-owned instances**, with full commit history rather than deployed artifacts.

| Tier | Location | Repos |
|---|---|---|
| App | `Scrivas_dev_env:/home/admin/` | `scrivas_backend` (1380 commits), `scrivas_gate` (183), `scrivas_search` (96), `patient` (59) |
| ML | `ML_dev:/srv/` | `post_processor` (137), `sai_suggestions` (93), `ml_monitoring` (62), `patient-summary-service` (49), `patient_document_parser` (41), `soniox_transcriber` (25) |

Every remote points at `git@git.devteam.space:scrivas/*` — the contractor's self-hosted GitLab, which Scrivas does **not** control. The on-instance checkouts are the client's only independent leverage over their own source.

**Gap:** both `/var/www` frontends are build output with no `.git`. Frontend source is not recoverable from EC2.

**Time-sensitive:** `scrivas_backend` received a commit on the assessment date. Development is active on infrastructure the contractor controls. Mirroring the repos to Scrivas-controlled storage is the recommended immediate action and is **not** included in this PR — it needs a scope decision first (it involves the contractor's GitLab credentials held on the dev box).

## Reliability triage

Production runs **23 containers on a single 15 GiB host** — including 3 Postgres instances, Kafka and OpenSearch — at **73% memory at rest**, with **no per-container memory limits** and **no swap** on any of the 7 instances. With no limits the OOM killer selects by resident size, so it would typically kill a database rather than the worker that caused the pressure. That matches the "random, unreproducible" symptom profile.

Kafka, OpenSearch and search-api additionally carry restart policy `no`, so a host reboot yields a partially-recovered stack that looks healthy from outside.

**Recorded as a structural exposure, not an observed root cause.** No OOM event appears in retained logs, `State.OOMKilled` is false on all 23 containers, and `RestartCount` is **0** on every one — the Celery workers showing "Up 6 hours" were redeployed, not crash-restarted. Confirming the hypothesis needs CloudWatch history these boxes do not retain, which is itself a finding and the basis for recommendations 3–5 in the report.

One encouraging contrast: the **ML tier is well-built** — ECR images tagged by commit SHA, blue/green slots, passing health checks. The application tier is compose-from-git-checkout with uncommitted `.env.save` and `docker-compose.yml.bkp` files in the prod working tree. The better pattern already exists in-house.

## Contents

| File | |
|---|---|
| `scripts/ec2_code_discovery.py` | EC2 inventory — describe + user data |
| `scripts/ec2_code_inspect.py` | read-only SSM probe set; commands reviewable in `PROBES` |
| `findings/ec2_code_discovery_report.md` | narrative writeup |
| `findings/code_dashboard.html` | client-facing dashboard, matching the existing design system |
| `findings/ec2_code_inspect*.json` | raw probe output |
| `index.html` | links the new dashboard and evidence |

## Review notes

- Probe output was **scanned for credentials before commit** ��� no AWS keys, passwords, tokens or private key blocks. The probe set reads manifests and VCS metadata, never file contents.
- Git metadata was read **as the owning user** (`sudo -u`) rather than by writing a `safe.directory` entry into root's gitconfig, to preserve the read-only guarantee.
- Several prod probes initially returned empty and were re-run with stderr visible; the empties were artifacts (`last` is not installed on prod), not clean health. Worth knowing when reading the JSON.
- The dashboard keeps the "what the evidence does not show" section in the **client-facing** version. If CloudWatch later points elsewhere, overclaiming a root cause would cost more credibility than the softer framing gains.

https://claude.ai/code/session_01YMxVaHXJsqpqKwncNQ9b1e
Co-authored-by: Alvaro Del Valle <alvaro.delvalle2@gmail.com>
Reviewed-on: #2
This commit was merged in pull request #2.
This commit is contained in:
2026-09-02 14:36:28 -06:00
co-authored by Alvaro Del Valle
parent 9e3d1f58d5
commit c987ed07c5
7 changed files with 1392 additions and 2 deletions
+11 -2
View File
@@ -55,10 +55,16 @@
<div class="eyebrow">Dasnuve · Cloud Discovery</div>
<h1>Scrivas — AWS Discovery Reports</h1>
<p class="lede">Read-only assessment of the Scrivas AWS Organization (<code>o-qfj0pvhhv7</code>) — footprint,
security posture, cost, and Organizations governance. Prepared to inform a proposal.</p>
<div class="meta">2 accounts · us-east-2 primary · generated 2026-08-19 · CONFIDENTIAL</div>
security posture, cost, and Organizations governance — extended with an EC2 source-code recovery and production reliability triage. Prepared to inform a proposal.</p>
<div class="meta">2 accounts · us-east-2 primary · generated 2026-08-19, EC2 code pass 2026-08-28 · CONFIDENTIAL</div>
<div class="cards">
<a class="card full" href="findings/code_dashboard.html">
<span class="k">CODE + RELIABILITY</span>
<h2>Source Code &amp; Reliability Triage <span class="badge">time-sensitive</span></h2>
<p>All 10 application repos found as full git checkouts on Scrivas-owned EC2 — recoverable despite the contractor holding the GitLab. Plus the reliability triage: 23 prod containers on one host, no memory limits, no swap.</p>
<span class="go">Open dashboard →</span>
</a>
<a class="card" href="findings/discovery_dashboard.html">
<span class="k">FINDINGS</span>
<h2>Discovery Findings</h2>
@@ -83,6 +89,9 @@
<h3>Written report &amp; raw evidence</h3>
<ul>
<li><a href="findings/discovery_report.md">discovery_report.md</a> — narrative writeup <code>(renders on GitHub)</code></li>
<li><a href="findings/ec2_code_discovery_report.md">ec2_code_discovery_report.md</a> — code recovery + reliability triage <code>(renders on GitHub)</code></li>
<li><code>findings/ec2_code_inspect.json</code> — SSM probe output, all 7 instances</li>
<li><code>findings/ec2_code_inspect_prod.json</code> — prod/stage probe snapshot</li>
<li><code>findings/org_assessment_report.json</code> — org / trust / delegated-admin raw data</li>
<li><code>findings/fast_discovery.json</code> — footprint, security, cost raw data</li>
<li><code>findings/member_lazka_547868853286.json</code> — member account raw data</li>