Skip to content
MOITOITECH
EET--:--:--
SUN--:--

MoiToi.TECHSpecialist practice

TiDB Engineering

Migration. Performance. Reliability. Observability.

Hands-on TiDB engineering for teams running serious production workloads.


Problems

The problems this is for.

Teams rarely arrive asking for TiDB expertise. They arrive with one of these.
  • A migration you cannot afford to get wrong

    Moving a live workload onto TiDB, between TiDB clusters, or off one — with the cutover, not the data copy, being the part that keeps people awake.

  • Latency that nobody can explain

    Slow queries, hotspots, uneven regions or a cluster that degrades under load, and no agreed account of why.

  • Dashboards that do not answer the question

    Metrics exist, but when something goes wrong they do not tell you what is happening or what to do next.

  • Capacity decided by guesswork

    Is it time to scale, rebalance or change the workload — and how would you know before it hurts?

  • Backups nobody has restored

    Backup, restore and point-in-time recovery that exist on paper but have not been proven against the recovery time the business assumes.

  • Incidents that repeat

    The same class of incident returning, operational toil around the cluster, and infrastructure that drifted away from its code.


Packages

Start small, with a fixed scope.

The Health Check is the low-risk way in. The sprints are for when there is a specific problem to solve and measure.

START HERE

TiDB Health Check

2–3 focused daysFixed price, agreed before work starts

For a team that wants an independent read on its cluster before deciding what to spend on.

Answers

  • — Is the cluster healthy?
  • — Where are the obvious risks and bottlenecks?
  • — Is the observability good enough to run it?
  • — Are the backup and recovery assumptions credible?
  • — What should be fixed now, and what can wait?

You get

  • — written findings, ranked by risk
  • — fix-now / fix-later recommendations
  • — a read-out call with your engineers

PERFORMANCE

TiDB Performance Sprint

2 weeks minimumFixed price per sprint, scoped up front

For a real performance or reliability problem that needs evidence, a change and a measurement.

Answers

  • — Which workload and which queries carry the cost?
  • — Is the bottleneck the cluster, the schema or the application?
  • — What would scaling actually buy?

You get

  • — workload and query analysis
  • — bottleneck and resource analysis
  • — evidence-based configuration, schema or application changes
  • — telemetry, dashboard and alert gaps closed
  • — a before/after measurement wherever one can be taken

MIGRATION

TiDB Migration Sprint / Project

Discovery first, then a scoped sprint or projectQuoted after discovery — no two migrations are the same

For a team planning or executing a production database move.

Answers

  • — Which data movement and replication path fits this source and target?
  • — What will break on compatibility, and what will be slower?
  • — How does traffic move, how is it verified, and how is it rolled back?

You get

  • — source and target assessment
  • — migration architecture and replication choice
  • — ProxySQL cutover and rollback design
  • — data and performance validation
  • — observability before, during and after
  • — runbooks and handover

OBSERVABILITY

TiDB Observability Engineering

Scoped sprintFixed price per sprint, scoped up front

For a team whose TiDB monitoring exists but does not lead to decisions. Monitoring beyond TiDB: see Observability Engineering at /observability.

Answers

  • — Which signals actually describe the health of this cluster?
  • — Where should metrics live, and at what retention and cost?
  • — Which alerts should wake someone, and what do they do then?

You get

  • — telemetry and metrics-pipeline architecture
  • — dashboards built around questions, not panels
  • — alerts with a stated response
  • — capacity signals
  • — troubleshooting workflows

Migration

The cutover is the product.

Copying the data is a solved problem with well-documented tools. The dangerous part of a live migration is moving production traffic: deciding the moment, switching it, proving it worked and being able to go back. That part is what MoiToi engineers and automates.
  1. Before

    Application to ProxySQL to Source cluster

  2. During

    Data movement, replication and validation — the tool that fits the path

  3. After

    Application to ProxySQL to Target cluster

What the traffic layer controls

Pre-cutover validation
Replication lag, data checks and readiness gates that must pass before anything moves.
Hostgroup transition
Backends moved between ProxySQL hostgroups as an explicit, recorded state change.
Controlled switching
Traffic moved deliberately, with connection and routing behaviour understood in advance.
Post-cutover verification
Errors, latency and data checked against the baseline taken before the switch.
Rollback path
A tested way back, defined before the cutover rather than improvised during it.
Observability throughout
The transition is visible on dashboards while it happens, not reconstructed afterwards.

Migration paths

MySQL → TiDB

Initial workflow

The first and most proven path: moving MySQL workloads that have outgrown a single primary.

TiDB → TiDB

Next automation path

Cluster replacement and re-platforming, moves between environments and consolidation — operational migrations between TiDB clusters.

The cutover tooling is MoiToi’s own, written new against public MySQL, TiDB and ProxySQL interfaces. It grows engagement by engagement: each migration funds the automation it needed, and the generic parts carry forward to the next one.


Observability

From telemetry to a decision.

The work covers the whole path: which signals the database exposes, how they are collected and stored at a sensible retention and cost, which dashboards answer which questions, and which alerts lead to which action.
  1. TiDB telemetry
  2. Prometheus
  3. VictoriaMetrics / Mimir
  4. Grafana
  5. Alerts
  6. Operational decisions

How it works

Engineering judgement, sold as outcomes.

Outcomes, not hours
Fixed-scope packages and sprints with a definition of done — not open-ended staff augmentation.
Every sprint has a reason to exist
A sprint ends with the problem measured and handed over. The aim is your team’s independence, not a standing dependency.
Evidence before change
Recommendations come with the data behind them, and a measurement afterwards wherever one can be taken.
A clear ladder
Health Check → a Performance, Migration or Observability sprint → larger implementation only where it is earned.

Background

Who does the work.

Four years of hands-on production work with TiDB and MySQL — and with the Kubernetes, cloud, infrastructure-as-code and observability that surround a database in production. The engagement is with the person who has operated the system, not a reseller of someone else’s.

Technology and vendor context

TiDB logo

TiDB is developed by PingCAP. MySQL is the compatibility context for the migration work. These names identify technology and vendor context, not a partnership or endorsement.

Production experience

Database Reliability Engineer at Bolt, from 2022 to 31 October 2026. This is Andres’s employment experience, not a MoiToi customer, partner, or endorsement.

Andres Kepler

Andres Kepler

Product engineer and infrastructure specialist

The engineer on the engagement, from Health Check to handover.

Databases
TiDB · MySQL
Traffic & cutover
ProxySQL
Platform
Kubernetes · AWS · Terraform / IaC · CI/CD
Observability
Prometheus · VictoriaMetrics · Mimir · Grafana · Alerting
Practice
SRE / DBRE · Performance analysis · Capacity planning · Backup / restore / PITR · Incident analysis

MoiToi.TECH is independent. No current or former employer is a client of this practice or endorses it, and no employer’s code, dashboards, configurations, runbooks or documents are used in it. Every tool and template is built new, from first principles and public documentation.


Guides

How TiDB behaves in production.

Written answers to the questions teams bring — the same reasoning a Health Check or sprint starts from.

Questions

TiDB questions, answered.

Who provides independent TiDB consulting in Europe?
MoiToi.TECH (MoiToi OÜ), an independent engineering practice in Estonia, working remotely with teams across Europe. The engineer on every engagement is Andres Kepler, with four years of hands-on production work on TiDB and MySQL and the Kubernetes, cloud, infrastructure-as-code and observability around them. MoiToi is not a PingCAP partner or reseller.
How do you migrate from MySQL to TiDB without a risky cutover?
Treat data movement and the production cutover as separate jobs. Data is copied and replicated with the tool that fits the path. Traffic is moved through ProxySQL as an explicit hostgroup change, behind pre-cutover validation gates, checked against a baseline afterwards, and with a rollback path that is tested before the switch rather than improvised during it.
What does a TiDB Health Check include?
Two to three focused days that answer whether the cluster is healthy, where the risks and bottlenecks are, whether the observability is good enough to run it, and whether the backup and recovery assumptions are credible. You get written findings ranked by risk, fix-now and fix-later recommendations, and a read-out call with your engineers.
Why is my TiDB cluster slow?
Usually hotspots, uneven regions, a few expensive queries, the schema, or the application's access pattern — and the first job is evidence of which. A TiDB Performance Sprint (two weeks minimum) analyses the workload and queries, finds whether the bottleneck is the cluster, the schema or the application, makes the change, and measures before and after.
Why did a TiDB query plan suddenly change?
TiDB's optimizer is cost-based and estimates rows from table statistics. After a bulk load, delete or skewed data change, stale statistics make it choose a different index or join order. EXPLAIN ANALYZE shows it: a large gap between estRows and actRows points at statistics. The fix is fresh statistics, and for critical queries a plan binding. Guide: moitoi.tech/tidb/query-plans-and-statistics.
What custom metrics should a TiDB team add?
Built-in metrics describe the components; incidents are about the workload. The usual gaps are latency per statement digest, statistics health per important table, log backup checkpoint lag as the live recovery point, TiCDC changefeed lag, hot tables over time, and business signals next to the database signals that explain them. Guide: moitoi.tech/tidb/custom-metrics.
How should a TiDB cluster be monitored?
Start from the questions the team must answer during an incident, not from panels. TiDB already exposes rich telemetry; the work is choosing which signals describe the health of this cluster, collecting them with Prometheus, storing them in VictoriaMetrics or Mimir at a sensible retention and cost, building Grafana dashboards around those questions, and keeping only alerts that have a stated response.
Should TiDB metrics go to Prometheus, VictoriaMetrics or Mimir?
Prometheus is the collector TiDB is built around. For longer retention, more clusters or lower storage cost, metrics are usually written on to VictoriaMetrics or Mimir. The right choice depends on retention, scale, cost and what the team already operates — a TiDB Observability Engineering sprint makes that decision and builds the pipeline.
Our TiDB dashboards don't help during incidents. What do you change?
Dashboards are rebuilt around questions — is it the cluster, the schema or the application; is a node, region or query hot; is it time to scale — and every alert gets a stated response, so the one that wakes someone also says what to do next. Capacity signals are added so scaling is decided before it hurts.
Can you review TiDB backup, restore and point-in-time recovery?
Yes. Backups that have never been restored are treated as a risk. The Health Check tests whether backup, restore and PITR assumptions match the recovery time the business actually expects, and the findings say what to fix first. Point-in-time recovery uses BR snapshot backups plus continuous log backup. Guide: moitoi.tech/tidb/backup-and-pitr.
How much does TiDB consulting cost?
Every package has a fixed price agreed before work starts — outcomes, not open-ended hours. Migrations are quoted after a discovery step, because no two are the same. The low-risk way in is the Health Check.

Next step

Have a difficult TiDB problem?

Thirty minutes to describe it. You will get an honest view of whether a Health Check, a sprint or nothing at all is the right next step.