Software · engineering management

Engineering analytics: delivery flow and DORA metrics from GitLab and tracker data, read-only and audited

An internal analytics system that reads a company's GitLab and issue tracker through a read-only connector and shows engineering management the delivery flow, DORA metrics and process quality, with MFA and an append-only audit log.

Industry
Software · engineering management
Technologies
Python 3.12 · FastAPI · SQLAlchemy 2 + Alembic
MFA

Challenge

Engineering management wanted a factual view of development activity, delivery flow and process quality across a large self-hosted GitLab and an issue tracker. Two constraints shaped the whole system. It had to be unable to change anything in GitLab, whatever token someone configured. And it had to support management decisions without turning into automatic scoring of people, so every number needed its context, its sample size and a clear note when data was missing.

Solution

We built one connector that every GitLab call goes through. It denies by default, allows only GET and HEAD requests, accepts GraphQL only as fixed query documents, protects against SSRF and redirects, and refuses admin, write-scope and maintainer tokens before they are stored. Imports are idempotent upserts, with daily aggregates recomputed after each run. Commit authors are matched to people through limited queries and push events, and about 98% of commits are attributed. Code computes every metric; the language model only puts computed aggregates into words.

How it connects

  1. 01 Read-only ingest
  2. 02 Identity resolution
  3. 03 Normalize & aggregate
  4. 04 Metrics & DORA
  5. 05 Dashboard & questions

What we delivered

  • Import of groups, projects, commits, merge requests, comments and pipelines
  • Identity resolution from commit email to account
  • Period comparison that accounts for data coverage and working days, with the sample size shown on every ratio
  • DORA delivery metrics: deploy frequency, lead time, change failure rate and rework; a DORA level is shown only when every input is available
  • Issue tracker link: tasks connected to merge requests, incidents, time to restore and the share of unplanned work
  • Roles and absences from an HR system's API or a CSV export
  • Dashboard with team, project and unmatched-author views, plus exports
  • An MCP server and an 'ask the data' panel
  • About 630 automated tests

Functionality

  • Default-deny connector: GET and HEAD only, fixed GraphQL documents, mutations rejected
  • Token check before saving: admin, write-scope and maintainer tokens are refused
  • Mandatory TOTP multi-factor sign-in, with no public sign-up
  • Append-only audit log, enforced in PostgreSQL
  • Secrets encrypted with Fernet and redacted from logs
  • Versioned prompts; only aggregates are sent to the model

Integrations

  • GitLab REST and GraphQL APIs, read-only
  • Issue tracker API
  • HR system API or CSV export
  • Google Gemini API

Technologies

  • Python 3.12
  • FastAPI
  • SQLAlchemy 2 (async) + Alembic
  • PostgreSQL 16
  • Redis + ARQ
  • httpx
  • React 18 + TypeScript + Vite
  • pytest

Result

Running on the company's real data, with a full year imported. Syncs are started on demand.

Why does the model only put numbers into words?

Managers act on these figures, so every figure has to be reproducible. Code computes each metric from stored data, with its period, coverage and sample size. The language model receives only those aggregates, under a versioned prompt, and turns them into a readable summary or an answer to a question. It never sees raw commits or comments and never produces a number of its own. If data for a period is incomplete, the dashboard says so instead of showing a confident figure.

Discuss your project

Tell us about your project