ai-agents-metrics

Architecture

What this document is: A technical map of the codebase — modules, data flow, storage, and integrations.

When to read this:

Related docs:


Summary

ai-agents-metrics is a CLI tool for analyzing AI agent work history, tracking spending, and optimizing workflows.

Primary layer — history pipeline: reads raw session files from ~/.codex or ~/.claude, extracts session structure, token cost, and session timelines, and stores results in a local SQLite warehouse. No prior instrumentation required.

There is no database server, no background process, and no network dependency.

Data flow:

CLI entrypoint (cli.py + cli_parsers.py + cli_constants.py)
  ↓  parses args and dispatches handlers through runtime_facade/
Commands (commands/ package)
  ↓  orchestration against narrow, command-specific runtime protocols
History pipeline (history/*)           ← primary analysis layer
  ↓  ingest → normalize → derive from ~/.codex or ~/.claude
  ↓  SQLite warehouse: session structure, token cost, session timeline
Reporting (report/)
  ↓  aggregates warehouse signals and renders a self-contained HTML report

How to read this


Directory Layout

ai-agents-metrics/
├── src/ai_agents_metrics/   # Main Python package
├── tests/               # Pytest test suite (grouped by subject area: cli/,
│                        # domain/, history/, reporting/, workflow/, infra/;
│                        # hypothesis strategies in tests/strategies/)
├── scripts/             # Automation and utility scripts
├── config/              # Public boundary rules (TOML)
├── pricing/             # Token pricing data
├── .githooks/           # commit-msg, pre-commit, pre-push hooks
└── pyproject.toml       # Package config, ruff, mypy, pytest settings

Package: src/ai_agents_metrics/

Entry Points

File / Package Role
__init__.py Version resolution: git-derived (commit_count.sha) with fallback to package metadata
__main__.py Enables python -m ai_agents_metrics dispatch
cli.py CLI dispatcher + facade surface for scripts/metrics_cli.py — records invocation, routes args.command to handlers, exposes console_main
cli_parsers.py Argparse parser construction (build_parser, per-group _add_*_parsers helpers, hidden-command filter)
cli_constants.py Path defaults (METRICS_JSON_PATH, CODEX_STATE_PATH, CLAUDE_ROOT, RAW_WAREHOUSE_PATH, …) consumed by both cli.py and cli_parsers.py
commands/ Thin CLI handlers grouped into history, report, install, and misc; _runtime.py defines narrow runtime protocols for individual handlers and composed workflows
runtime_facade/ Concrete runtime composition surface for history orchestration, reporting, pricing, audits, and installation

runtime_facade/contracts.py statically verifies that the concrete facade satisfies every command-specific runtime protocol. The module is checked by mypy and has no runtime behavior.

Core Domain

File Role
domain/ Domain package split into submodules: models.py (dataclasses), serde.py (from_dict / to_dict — the only place that converts timestamps between str and datetime), validation.py, aggregation.py, ids.py, time_utils.py. Public API re-exported via domain/__init__.py.
storage.py Atomic file writes and fcntl lockfile helpers

History Pipeline

Sequential stages that reconstruct goal history from raw Codex + Claude Code agent state:

Codex (~/.codex) or Claude Code (~/.claude)
  ↓  history/ingest/         → .ai-agents-metrics/warehouse.db (raw_* tables)
       ingest/warehouse.py     — schema, SQL helpers, manifest, path resolution
       ingest/codex.py         — Codex adapter (state_5.sqlite, logs_1.sqlite, session JSONL)
       ingest/claude.py        — Claude Code adapter (projects/*.jsonl, subagent files)
       ingest/__init__.py      — orchestrator (ingest_codex_history), IngestSummary, snapshots
  ↓  history/normalize.py    → cleaned warehouse rows (normalized_* tables)
  ↓  history/classify.py     → session kinds (main vs subagent) + practice-event labels
  ↓  history/derive.py       → warehouse analysis marts (derived_* tables)
       history/derive_build.py    — pipeline stage builders
       history/derive_insert.py   — typed inserts into derived_*
       history/derive_schema.py   — schema for derived tables
       history/project_paths.py   — canonical checkout and worktree path rules

For the layering rules (raw_* byte-perfect, normalized_* typed, derived_* aggregated) see warehouse-layering.md.

Analysis and Reporting

File Role
report/html_report.py Public facade for the HTML report: re-exports aggregate_report_data and render_html_report
report/application/ Typed report contracts, query and pricing ports, and HTML-report use cases; application code belongs in this package and receives no persistence rows
report/sqlite_query.py SQLite adapter that owns SQL, schema knowledge, and persistence-row mapping behind the query port
report/aggregation.py Transforms warehouse retry, token, model, and practice rows into chart-ready series
report/buckets.py Pure date/time-bucket helpers (parse, bucket key, make buckets)
report/template.py Self-contained HTML/CSS/JS template string; no Python logic
warehouse/application.py Shared typed readiness, scope, and token-breakdown application contracts
warehouse/domain.py Pure token-breakdown value objects, aggregation, ranking, and remainder rules
warehouse/adapters/ Concrete SQLite warehouse adapters behind application ports
runtime_facade/breakdown.py Composition root for the breakdown use case and its concrete adapters
runtime_facade/summary.py Composition root for warehouse summary loading and its readiness gate

Concrete persistence adapters are imported only by runtime_facade composition modules. Adapter package internals may collaborate with modules inside the same adapter package; application, presentation, and command modules depend on typed ports instead.

Integrations

File Role
usage/pricing_runtime.py Sanctioned application-level pricing API: resolves effective pricing path and loads workspace-aware pricing
usage/resolution.py Pricing data loading, usage event parsing, cost computation, and window resolution logic for Claude and Codex sessions
usage/backends.py UsageBackend Protocol with ClaudeUsageBackend and UnknownUsageBackend implementations; delegates window resolution to usage.resolution
git_hooks.py Implements commit-msg validation and pre-push security scanning logic
commit_message.py Validates commit subject format (CODEX-123: / NO-TASK:)
public_boundary.py Verifies files against TOML-configured inclusion/exclusion rules
completion.py Shell tab-completion helpers

Data and Storage

Primary store: .ai-agents-metrics/warehouse.db

The warehouse is disposable derived state, not a durable compatibility boundary. Schema changes do not require backward-compatible readers or migrations for existing local databases. Commands must reject an outdated schema clearly; running history-update rebuilds the warehouse from the source agent history.


CLI Entry Points

Installed command: ai-agents-metrics → ai_agents_metrics.cli:console_main

Boundary note:

Key command groups:

Group Commands
Inspection show, render-html
History pipeline history-ingest, history-normalize, history-classify, history-derive, history-update
Tooling install-self, completion, verify-public-boundary, security

Scripts (scripts/)

Script Purpose
metrics_cli.py CLI entry point shim for local development
public_overlay.py Bidirectional sync between private repo and oss/ public mirror
build_standalone.py Builds self-contained binary distribution

Tests (tests/)

One test file per module; files are grouped into subject-area subdirectories so the root shows structure at a glance:

Subdir Test file Covers
cli/ test_metrics_cli.py Full CLI workflow integration
domain/ test_metrics_domain{,_properties}.py Domain model logic + hypothesis invariants
history/ test_history_{ingest,normalize,normalize_properties,derive,classify,compare,audit,pipeline_json}.py Pipeline stages
reporting/ test_{html_report,report_application,sqlite_report_query,show_json}.py Pure rendering, report use cases, SQLite report adapter, and summary JSON
workflow/ test_commit_message.py Commit-message and hook integration
infra/ test_{public_boundary,public_overlay,security}.py Boundary and security rules
strategies/ domain.py, history.py Hypothesis strategies shared across property tests
tests/private/ (private root only) test_git_hooks.py, test_claude_md.py Git hook behavior and doc generation

conftest.py provides shared fixtures (temp metrics paths, fake goal factories, etc.) plus find_repo_paths() — a [tool.codex_tests]-marker-based helper for resolving the repo root from any test subdir. Prefer it over Path(__file__).parents[N] so test paths stay stable when files move between subdirs.


Code Quality Configuration

Tool Config Settings
ruff pyproject.toml 15 rule categories (B, C4, ERA, F, FURB, I, PERF, PGH, PTH, Q, RET, RSE, SIM, TC, UP); target Python 3.14; line length 100
mypy pyproject.toml strict = true at top level (ARCH-030); covers src/ and scripts/
import-linter pyproject.toml Executable dependency contracts for domain, storage, history, reporting, command, usage, and CLI boundaries
Semgrep CE config/semgrep/ Project-specific semantic guards against persistence leakage into controllers and direct I/O in application modules
pylint pyproject.toml + Makefile Full default rule set on the whole project (ARCH-019 … ARCH-023); a small disable list documents the few intentionally-off rules
hypothesis pyproject.toml dev dep Property-based tests for domain/aggregation (8 invariants) and history/normalize (8 invariants); strategies in tests/strategies/
pytest pyproject.toml pythonpath = ["src"], xdist auto workers, 5s default timeout (overridden per-test on hypothesis suites)
coverage pyproject.toml Branch coverage, parallel mode, source = ai_agents_metrics

Makefile targets: lint, security, typecheck, test, arch-check, verify, verify-fast, coverage, package, public-overlay ops.

Git hooks (.githooks/):