MSP & multi-tenantadvanced7 min read

Multi-tenant monitoring for MSPs: separating forty customers without forty installations

One Zabbix per customer doesn't scale; one Zabbix for everyone leaks. This is the architecture that does work with open-source building blocks — Prometheus per customer, Grafana organisations, Alertmanager routing on a tenant label, one naming scheme and one offboarding procedure — and the honest bill for what it costs in maintenance at ten, forty and a hundred customers.

Contents
  1. Step 0: the five requirements that make an MSP different
  2. Step 1: naming and tagging — the foundation everyone skips
  3. Step 2: collection — one Prometheus per customer, one central Thanos or Mimir
  4. Step 3: dashboards — Grafana organisations, not folders
  5. Step 4: alerts — one rule set, routing on tenant
  6. Step 5: customer reports and evidence
  7. Step 6: offboarding — one script, demonstrable
  8. The bill
  9. Pitfalls
  10. What you still don't have
  11. How monsys does it
  12. FAQ

An MSP with forty customers has two bad options and one hard one. Bad: forty separate monitoring installations, each with its own version, its own alert rules and its own password that one colleague knows. Also bad: one big installation in which customer A can in principle see customer B's hosts if someone forgets a filter. Hard: one platform with real separation — per customer its own data, its own access, its own alerts, its own reports — and still one way of working. This article builds that third option with open-source building blocks, and is honest about where it hurts.

Step 0: the five requirements that make an MSP different

Before installing anything, the list every choice gets checked against:

RequirementWhyWhat it means technically
IsolationCustomer A must never see anything of customer B, not even by accidentSeparation at the data level, not just in the UI
Per-customer accessCustomer A's IT lead wants to look themselvesRead-only accounts per customer, without you managing an installation per customer
One way of workingYour team must be able to take over any customerThe same alert rules, the same naming, the same runbooks — everywhere
Evidence per customerNIS2 customers ask for reports; you're their supplier (Art. 21(2)(d))Exportable, per customer, with a date
OffboardingContracts end; GDPR requires the data to be gone thenOne procedure that removes everything of customer X, demonstrably

Step 1: naming and tagging — the foundation everyone skips

Everything that follows hangs on one label: the customer. Choose the format now and never deviate from it.

# Every host, every metric, every alert gets these labels:
tenant: acme            # short, stable customer code; never the company name (that changes)
env: prod               # prod | staging | dev
role: web               # web | db | app | edge | backup
site: brussels-dc1      # or cloud region

On the host itself the customer code lives in one place, and every tool reads it there:

sudo tee /etc/monitoring-identity >/dev/null <<'EOF'
TENANT=acme
ENV=prod
ROLE=web
SITE=brussels-dc1
EOF
# Hostname follows the convention: <tenant>-<env>-<role>-<nn>
hostnamectl set-hostname acme-prod-web-01

Sounds bureaucratic. But it's the difference between {tenant="acme"} in every query and a spreadsheet of "which IP belongs to whom".

Step 2: collection — one Prometheus per customer, one central Thanos or Mimir

The simplest hard separation: every customer gets their own Prometheus (on a small VM at your place, or on-prem at the customer), and those Prometheuses write through to one central store with the tenant label as key.

# /etc/prometheus/prometheus.yml on customer acme's Prometheus
global:
  external_labels:
    tenant: acme              # attached to EVERY metric before it leaves the building
scrape_configs:
  - job_name: node
    file_sd_configs:
      - files: ['/etc/prometheus/targets/*.yml']   # acme's hosts, managed via Ansible
remote_write:
  - url: https://metrics.msp.example.be/api/v1/push
    basic_auth:
      username: acme
      password_file: /etc/prometheus/remote-write.pass
    headers:
      X-Scope-OrgID: acme     # Mimir/Cortex tenant header; Thanos uses the external_label

Centrally runs Grafana Mimir (or Thanos Receive) with multi-tenancy on: the X-Scope-OrgID determines in which physically separate set of blocks the data lands. A query without a tenant header gets nothing. That's isolation at the data level, not the UI level.

The price: a Prometheus VM per customer (1 vCPU, 2 GB is enough up to ~50 hosts), plus a central Mimir cluster you have to understand. With ten customers that's an afternoon per month; with forty, half an FTE.

Step 3: dashboards — Grafana organisations, not folders

Grafana has teams and folders (separation in the UI) and organisations (separation of datasources, users and dashboards). For customer access you use organisations. A user in org "acme" sees no datasource from org "globex", full stop.

# Per customer an org, with a datasource that passes the tenant header
curl -s -u admin:$GRAFANA_ADMIN_PW -X POST https://grafana.msp.example.be/api/orgs \
  -H 'Content-Type: application/json' -d '{"name":"acme"}'
ORG_ID=$(curl -s -u admin:$GRAFANA_ADMIN_PW https://grafana.msp.example.be/api/orgs/name/acme | jq .id)
curl -s -u admin:$GRAFANA_ADMIN_PW -X POST https://grafana.msp.example.be/api/datasources \
  -H "X-Grafana-Org-Id: $ORG_ID" -H 'Content-Type: application/json' -d '{
    "name":"metrics","type":"prometheus","url":"https://metrics.msp.example.be/prometheus",
    "access":"proxy","jsonData":{"httpHeaderName1":"X-Scope-OrgID"},"secureJsonData":{"httpHeaderValue1":"acme"}}'

The dashboards themselves you manage as code (JSON in git, provisioning per org) so every customer gets the same "Fleet overview" dashboard and a fix lands everywhere at once. Your own team gets a separate "MSP" org with a datasource without tenant filter (Mimir: X-Scope-OrgID: acme|globex|... for cross-tenant queries) — that's the only place where everything is visible together, and that org has MFA mandatory.

Step 4: alerts — one rule set, routing on tenant

Alert rules you write once and provision to every customer Prometheus. Alertmanager centrally routes on the tenant label to the right channel, and to the right on-call.

# alertmanager.yml (central)
route:
  receiver: msp-default
  group_by: ['tenant', 'alertname']
  routes:
    - matchers: ['tenant="acme"']
      receiver: acme
      continue: true                 # both to the customer AND to us
    - matchers: ['tenant="globex"']
      receiver: globex
      continue: true
    - matchers: ['severity="critical"']
      receiver: msp-oncall
receivers:
  - name: msp-default
    webhook_configs: [{ url: 'https://ntfy.msp.example.be/all' }]
  - name: msp-oncall
    webhook_configs: [{ url: 'https://ntfy.msp.example.be/oncall' }]
  - name: acme
    email_configs: [{ to: 'it@acme.example', from: 'monitoring@msp.example.be', smarthost: 'mail.msp.example.be:587' }]
  - name: globex
    webhook_configs: [{ url: 'https://ntfy.msp.example.be/globex' }]

The trap: a receiver pointing to a shared channel with several customers in it. One Teams channel with "all customer alerts" is a data breach in the making. Per customer its own channel or email address, always.

Step 5: customer reports and evidence

A NIS2 customer asks you to demonstrate that their servers are patched, monitored and backed up — you're their supplier under Art. 21(2)(d). That report has to be per customer, dated, and without anything from another customer in it.

# Monthly per tenant: availability, open alerts, hosts, patch status from Prometheus
TENANT=acme; MONTH=$(date -d 'last month' +%Y-%m)
Q='https://metrics.msp.example.be/prometheus/api/v1/query'
H="X-Scope-OrgID: $TENANT"
{
  echo "# Report $TENANT — $MONTH — $(date -Is)"
  echo "hosts: $(curl -s -H "$H" "$Q" --data-urlencode 'query=count(up{job="node"})' | jq -r '.data.result[0].value[1]')"
  echo "uptime%: $(curl -s -H "$H" "$Q" --data-urlencode 'query=avg_over_time(up{job="node"}[30d])*100' | jq -r '.data.result[0].value[1]')"
  echo "reboot-required: $(curl -s -H "$H" "$Q" --data-urlencode 'query=count(node_reboot_required == 1)' | jq -r '.data.result[0].value[1] // 0')"
  echo "critical alerts (30d): $(curl -s -H "$H" "$Q" --data-urlencode 'query=count_over_time(ALERTS{severity="critical",alertstate="firing"}[30d])' | jq -r '[.data.result[].value[1]|tonumber]|add // 0')"
} > "/srv/reports/$TENANT/$MONTH.txt"
sha256sum "/srv/reports/$TENANT/$MONTH.txt" >> "/srv/reports/$TENANT/MANIFEST"

Sign the manifest file as in the evidence guide; then the customer (or their auditor) can verify the report hasn't been altered.

Step 6: offboarding — one script, demonstrable

When the contract ends, everything of that customer has to go: metrics, dashboards, users, alert routes, and the hosts from your inventory. Write the script before you onboard the first customer.

#!/usr/bin/env bash
# offboard.sh <tenant> — removes everything of a customer, logs every step
set -euo pipefail
T=$1; LOG=/srv/offboarding/$T-$(date +%F).log
log(){ echo "$(date -Is) $*" | tee -a "$LOG"; }
log "start offboarding $T"
# 1. Grafana org (takes dashboards, datasources and org users with it)
ORG=$(curl -sf -u admin:$GRAFANA_ADMIN_PW https://grafana.msp.example.be/api/orgs/name/$T | jq .id) && \
  curl -sf -u admin:$GRAFANA_ADMIN_PW -X DELETE https://grafana.msp.example.be/api/orgs/$ORG >/dev/null && log "grafana org $ORG deleted"
# 2. Metrics: Mimir tenant deletion (asynchronous; marks all blocks for deletion)
curl -sf -X POST -H "X-Scope-OrgID: $T" https://metrics.msp.example.be/compactor/delete_tenant && log "mimir tenant deletion requested"
# 3. Alertmanager route + Prometheus VM
sed -i "/tenant=\"$T\"/,+2d" /etc/alertmanager/alertmanager.yml && systemctl reload alertmanager && log "alert route removed"
log "TODO manual: remove Prometheus VM $T-prom; revoke remote-write credentials; hosts out of Ansible inventory"
# 4. Reports: keep per contract (usually 1-3 years), then delete — scheduled separately
log "reports retained until $(date -d '+3 years' +%F) per contract"
log "done"

The log file is the evidence you give the customer: "on date X everything was removed, except the reports we contractually retain until Y".

The bill

10 customers40 customers100 customers
Prometheus VMs1040100
Maintenance (updates, disk space, rolling out rules)~4 h/month~20 h/month~1 FTE
Mimir/Thanos cluster1 node suffices3 nodes, someone who understands itdedicated
Grafana orgs + provisioningscriptscript + testsscript + tests + an owner
Risk of tenant leak through human errorlowmedium (every new colleague)real

The model works. At forty customers it costs half a person to keep working, and that person has to genuinely know Prometheus, Mimir, Grafana provisioning and Alertmanager routing.

Pitfalls

  • Label leak via external_labels. Forget tenant on one Prometheus and those metrics land without a tenant and appear — depending on your Mimir config — at the default tenant or nowhere. Test every new customer Prometheus with a query before taking it into production.
  • Shared agent tokens. One node_exporter basic auth for all customers means a compromised host of customer A can push (or read) metrics of customer B. Per tenant its own credentials, and rotation at offboarding.
  • Grafana admin who can get in everywhere. Handy, until a colleague leaves. The MSP org with cross-tenant access has MFA, its own audit log and a quarterly review of who's in it.
  • Customer name in the hostname. acme is a code. "Acme Industries NV" changes at a merger and then sits in 400 dashboards. Codes don't change.
  • Forgetting the data processing agreement. You monitor their systems: that's processing of personal data (IPs, usernames in logs). Every customer signs a DPA, and the offboarding procedure refers to it.

What you still don't have

  • Everything that isn't a metric. Prometheus is for numbers. CVEs per customer, inventory, who-has-sudo, backup freshness, certificates: those are separate pipelines with separate tenant separation — and you have to rebuild each of them with the same five requirements.
  • Intervening. "Restart that service at customer A" is an ssh session, with the question who may do that, who did it, and whether customer A may see it. A signed trail of that doesn't exist.
  • Evidence that the separation works. A customer asking "how do I know my competitor can't see my data?" gets an architecture diagram. An auditor wants more.

How monsys does it

monsys is multi-tenant from the first migration: every table has a tenant_id, every query pins it explicitly and PostgreSQL row-level security is the backstop. One hub, forty tenants, zero extra VMs. Per tenant: its own agents with their own tokens and mTLS certificates, its own users with scope-based roles (viewer on customer A, editor on customer B), its own branding and its own ntfy topic. For you as MSP: a cross-customer overview over all tenants where you're staff, a handover report per customer (Ed25519-signed, offline-verifiable) and on-call rotations per group. Intervening goes via Emergency Action Tokens logged per tenant, so customer A sees in their own audit log what you did on their hosts — and nothing of customer B.

FAQ

Can't I just use one Prometheus with a tenant label?

Technically yes, but then the separation is a filter in the UI and in every query. One forgotten {tenant="acme"} in a dashboard variable and customer A sees customer B. For an MSP under NIS2 that's an unacceptable risk; the separation has to be at the storage level.

How many customers can I handle before I need this?

Up to five customers, "a small installation per customer" works fine. From ten, maintaining separate installations becomes more expensive than a central setup. From forty, the central setup itself is half a job.

What does a customer under NIS2 concretely ask of their MSP?

Demonstrability: that their systems are monitored, patched and backed up, with dates and without data from other customers. Plus a contract (DPA and a security annex) and an incident procedure: if you see a break-in at customer A, customer A must be able to notify the CCB within 24 hours, so you have to tell them even faster.

Written by the monsys team — sysadmins who do this every day.

Done it by hand? Let monsys keep it running.

Everything in this guide runs in monsys as a continuous check, with history, alerts and audit evidence. 5 servers free, EU-hosted in Belgium, installed in 60 seconds.