DBA Brain

A database operations toolkit for people who are on call.

pip install dbabrain · PyPI · Downloads · GitHub

What it does

DBA Brain collects health metrics from your SQL Server, Oracle, PostgreSQL and MySQL instances, grades them against policies you write, runs your scheduled SQL, takes backups and proves them by restoring them, checks objectives against real measurement history, and tells you about it, on a schedule, without an agent on any monitored machine.

It sends nothing anywhere by itself. No telemetry, no usage reporting, no update check. It connects to the databases and hosts you list, and, only if you turn it on, to a chat service. Nothing else.

Health metrics

Around ninety metrics across four engines: availability, capacity, performance, recoverability, security, maintenance. Collected and normalised into one shape.

Graded, not dumped

Every measurement is graded against a policy you wrote. You read a verdict (OK, WARNING, CRITICAL) and pull the evidence only for what came back bad.

Backups proven by restore

Runs backups, restores them onto a disposable target, verifies the result and records what was proven and when. The restore is the proof.

Scheduled SQL

Your own SQL, on a schedule or on request, against approved targets, delivered as text or a spreadsheet. The SQL is a reviewable file, never a string in config.

Reports

Fleet inventory, per-server metrics, index usage and SLA pages, with a freshness gate so a stale number is never reported as a current one.

SLA / SLO

Indicators computed from stored measurement history, objectives evaluated with an error budget, and how much of it is left.

Alerts and chat commands

Telegram delivery, one message at a time. Commands people send back are gated by the person's clearance and the chat's.

Config as code

Every decision is a JSON file you keep in version control. A threshold change is a diff and a review, not an edit to somebody's script.

Secrets kept encrypted

PBKDF2-HMAC-SHA256 + Fernet. The passphrase is supplied at run time and never written to disk, a log or the store. Passwords travel on stdin, never argv.

What it is not

What it produces

Real report pages from a live estate, with every name replaced by a stable fake one. Click an image to open it full size.

Fleet inventory and health report: server counts per engine, critical findings ordered by severity with a recommended action
Fleet inventory: every instance, its health, and a priority list of what to act on first, each with a recommended next step.
Per-server metrics page: health score, status cards for availability, CPU, memory, disk, backup and security, and a list of problems
Server metrics: one card per health area, showing the worst reading, the rule that graded it and what it means.
SLA / SLO compliance page with overall verdict and per-engine cards
SLA / SLO compliance: verdict per engine, computed from stored history with no database connections.
Index usage report for one SQL Server instance: total, unused, cold and droppable indexes
Index usage: total, unused, cold and droppable indexes per instance, with how long the counters cover.

How it works

One run, end to end

Whoever starts it (the daemon, a person, a chat command or an AI agent), a run takes the same path. The app hands one JSON request to transport, which starts the shared operations CLI; that CLI reaches the database or host and answers with one JSON envelope. Results land in the runtime store, and reports and alerts are built from there.

CALLERS APP SHARED LAYERS YOUR ESTATE Daemon (schedule) Person (shell) Chat command AI agent App CLI metrics sql_tasks backup_restore sla · reports sre · control one JSON object in, one JSON object out transport starts process common.cli JSON on stdin SQL Server Oracle PostgreSQL MySQL / MariaDB Linux / Windows read-only login, SSH or WinRM job_runs, measurements Runtime store SQLite or PostgreSQL Reports Web host :8080 Telegram alerts

Every command answers in the same envelope, whether the work succeeded or not:

{
  "success": true,
  "operation": "restore-full",
  "message": "Restored SALESDB_STG to 2026-08-07 01:40:00.",
  "error": null,
  "data": { },
  "metrics": { "duration_ms": 41230 }
}

From a measurement to an alert

Collect~90 metrics per engine
→
Normaliseone shape for all engines
→
Gradeagainst your policy in data/*.json
→
Storeruntime store
→
Report / alert / SLAby severity and error budget

A backup is proven by restoring it

Backupfull / diff / log
→
Copyto a disposable target
→
Restorelatest or point in time
→
Verifythe result
→
Recordwhat was proven, and when
→
Clean upand apply retention

Why: cheap for an AI agent, fine without one

The command a person runs is byte-for-byte the command an agent runs, so there is one code path to audit, and every run is logged the same way whoever started it.

  WITHOUT a tool                        WITH DBA Brain
  agent -> raw SQL -> agent             agent -> DBA Brain -> database

  "check this database"                 "check this database"
    -> SELECT ... dm_os_wait_stats        -> one JSON request
    <- every row, into the context
    -> SELECT ... sys.databases           <- one graded envelope:
    <- every row, into the context             41 checks, 40 OK,
    -> ... once per check, per run             1 WARNING: appdb
    <- all of it, healthy or not               log backup 31h old

The saving is not compression. It is not sending what nobody needed to read. And no model is required: remove the agent and the daemon, schedules, reports and alerts all still run.

15 components

Ten apps, each doing one job with its own CLI and never importing another app, and five shared layers that every app stands on. The list is closed, and every component has exactly one reference doc.

Who may call whom

10 apps jobs · metrics · sql_tasks · reports · telegram backup_restore · sla · sre · control · webhost import transportthe one client dbruntime store logging_opslogs, job_runs commonoperations, CLI only starts common.cli lib pure rules: imported everywhere, runs nothing

Apps

03App command daemon

db_ops/jobs · db-ops daemon

The scheduler: runs each app on its own interval inside its allowed hours, skips one that is still running.

04Metrics engine

db_ops/metrics

Around ninety metrics across four engines, collected and normalised into one shape.

05SQL task runner

db_ops/sql_tasks

Your own SQL, on a schedule or on request, against approved targets, as text or a spreadsheet.

06Reports

db_ops/reports

Scheduled reports and inventory pages, with a freshness gate against stale numbers.

07Chat delivery and commands

db_ops/telegram

Delivers the outgoing queue and executes commands, gated by the person's and the chat's clearance.

08Backup / restore

db_ops/backup_restore

Runs backups, restores them onto a disposable target, verifies, and records what was proven.

09SLA / SLO compliance

db_ops/sla

Indicators from stored measurement history, objectives with an error budget.

10SRE

db_ops/sre

Provisions disposable lab databases for drills: single instances or small HA clusters, in Docker or on VMs.

11Control

db_ops/control

Builds and deploys the toolkit to another node, and watches the toolkit itself.

12Web host

db_ops/webhost

Serves the rendered reports over HTTP and hosts the console. Publishes files, never generates them.

Shared layers

01Runtime store

db_ops/db

The toolkit's own data: job runs, measurements, report state, delivery queue, restore history. SQLite to start, PostgreSQL later.

02Logging engine

db_ops/logging_ops

Scoped application logs, runtime logs, shared errors and daily archives.

13Common

db_ops/common · CLI only

Reaching a host, running SQL, moving a file, rotating a password. Every command takes one JSON object.

14Lib

db_ops/lib · import only

Pure rules: time windows, notify routing, severity, formatting. Imports nothing from the rest.

15Transport

db_ops/transport

The one client: starts common.cli and db.cli for every component. Imports only lib.

Install

Python 3.12+. Each database driver is an extra, so you install only what you run.

[postgres]PostgreSQL, pure Python
[mysql]MySQL / MariaDB, pure Python
[oracle]Oracle, no client library needed for 12.1 and newer
[mssql]SQL Server, also needs Microsoft's ODBC driver
[ssh], [winrm]OS-level metrics on Linux / Windows hosts
[all]everything
python -m venv .venv
.venv/bin/pip install 'dbabrain[postgres]'

Published on PyPI via trusted publishing.

Five-minute try-out with a throwaway PostgreSQL container:

cd examples/postgres-quickstart
docker compose up -d
python -m db_ops.db.cli      --config config.json init
python -m db_ops.metrics.cli --config config.json collect --dry-run
python -m db_ops.metrics.cli --config config.json collect
python -m db_ops.metrics.cli --config config.json report

Docs and source

Contact: donations and investment

DBA Brain is free, open source, and maintained by one person. If you would like to donate, sponsor the project, or talk about investment or partnership, please get in touch directly:

Emailtanthanhkaka01@dbabrain.dev
Emailtanthanhkaka01@gmail.com
Phone+84 888 783 789 (0888 783 789)

For bugs and feature requests, please use GitHub issues. For security problems, follow the security policy rather than email.