Data engineering & analytics
The full data path — from tracking and ingestion pipelines through to the dashboards and reports that turn raw events into decisions, designed by an engineer who has run real-time analytics at billions of events a month.
Last updated:Data engineering is the practice of building the pipelines and storage that move raw events into a structured, queryable form for analysis. Revenant Systems builds the full data path — tracking and ingestion through to warehousing, dashboards, and reporting — using Python, PostgreSQL, ClickHouse, DuckDB, and Airflow.
Who is this data engineering service for?
Teams whose numbers have become business-critical while the path producing them has not kept up: founders reconciling spreadsheets that disagree, operations leads running month-end by manual export, and data owners whose dashboards contradict each other. Where the need is only an off-the-shelf BI tool configured, that vendor's partners are the cheaper route — this service builds and repairs the path underneath such tools.
Which entry point fits your data problem?
Start at the layer that hurts. A slow production database starts with the PostgreSQL performance audit, from £4,950; an analytical cluster under strain starts with the ClickHouse equivalent, from £7,500; reporting trapped in a spreadsheet starts with the Excel workbook audit, from £3,950. Implementation that follows — pipelines, warehousing, reporting — is typically £15,000–£60,000, scoped against what the audit found.
| Engagement | Price |
|---|---|
| PostgreSQL Performance Audit | From £4,950 |
| ClickHouse Performance Audit | From £7,500 |
| Excel Workbook Audit | From £3,950 |
| Data engineering implementation | Typically £15,000–£60,000 |
UK company and contracting entity · UK GDPR-aware delivery · Senior in-house delivery · Code and documentation owned by you
What does the data path include?
The data path is the route raw events take from source to decision: capture, ingestion, storage, transformation, and reporting. A complete data path makes the same numbers reproducible across dashboards and reports, so teams can trust what they measure rather than reconciling conflicting figures.
Revenant Systems builds each stage — ingestion in Python and Airflow, storage in PostgreSQL, ClickHouse, or DuckDB, and reporting on top — as one pipeline rather than disconnected scripts, drawing on the founder's firsthand experience running ClickHouse through hundreds of billions of rows.
- Event tracking and ingestion
- Warehousing and data modelling
- Dashboards and reporting
- Scheduling and orchestration (Airflow)
Which databases does Revenant use, and when?
Revenant Systems selects the data store by workload. PostgreSQL handles transactional and general-purpose data; ClickHouse handles high-volume analytical queries; DuckDB handles fast in-process analytics and local transformation. Matching the engine to the workload keeps queries fast and infrastructure simple — and most estates need only PostgreSQL until analytical volume and cardinality outgrow it.
| Engine | Workload | Chosen when |
|---|---|---|
| PostgreSQL | Transactional and general purpose | The default store for application data |
| ClickHouse | Large-scale analytics | Analytical volume and cardinality outgrow the transactional store |
| DuckDB | In-process analytics | Fast local analysis and transformation without a server |
| Airflow | Orchestration | Pipelines need scheduling, dependencies, and visibility |
What's included
- Tracking and ingestion pipelines
- Data warehouse design
- Transformation and data modelling
- Dashboards and reporting
Stack Python · PostgreSQL · ClickHouse · DuckDB · Data Warehousing · Airflow
Frequently asked questions
Our dashboards and reports disagree — why?
Usually because the numbers travel different paths: separate scripts, exports, and hand-maintained queries that each transform the data slightly differently. When every dashboard and report draws from the same audited path, a figure has exactly one definition — and the weekly argument about whose spreadsheet is right goes away.
Can you build the event tracking as well as the warehouse?
Yes — the engagement covers the full path, from event tracking and ingestion through warehousing and data modelling to the dashboards and reports on top. Building the stages as parts of one design, rather than as scripts that happen to feed each other, is the point of the service.
What scale does this experience come from?
The founder's prior roles: real-time analytics platforms ingesting billions of events a month, and ClickHouse warehouses that have worked through hundreds of billions of high-cardinality rows. Smaller estates get the simplest pipeline that stays correct — scale experience informs the design; it doesn't inflate it.
Do we need ClickHouse, or will PostgreSQL do?
The split is clean: PostgreSQL for transactional and general-purpose data, ClickHouse where analytical query volume and cardinality outgrow it, DuckDB for fast in-process analytics. Which one you need is settled at the audit stage, from your own query load rather than from fashion.
Our existing PostgreSQL is slow — can you help?
Yes — that is a fixed-scope package: the PostgreSQL performance audit examines query load, indexing, autovacuum and bloat, configuration, and storage overhead, and delivers a severity-ranked report with a remediation roadmap. The same audit exists for ClickHouse. Either is a common starting point for wider data work.
Can you replace reporting that lives in a spreadsheet?
Yes, and the safe route starts with understanding it: the Excel workbook audit maps the workbook's actual behaviour before pipelines and dashboards replace it. Reporting logic that grew up in a spreadsheet usually contains business rules nobody has written down anywhere else.
Our reporting still runs on manual exports — is that fixable?
Yes — manual exports and copy-paste steps are usually the symptom of a missing pipeline stage, and they are where numbers quietly diverge. Replacing them with automated ingestion and transformation makes the reporting reproducible and removes the person-shaped single point of failure in the process.
Every engagement follows the same process — see how we work.
Have a data problem to solve? Let's talk.
Get in touchYour message is read by the engineer who would scope the work; the reply is a short technical conversation.