What is synthetic monitoring?
Scripted checks that simulate real users
Synthetic monitoring continuously runs scripted checks against your application from monitoring locations around the world. The scripts simulate real user interactions — loading a page, calling an API, logging in, completing a purchase — and record performance and availability.
The word "synthetic" distinguishes these scripted probes from real user monitoring (RUM), which measures what real users actually experience. Both are useful; neither replaces the other.
Synthetic checks run on a schedule (every 1, 5, or 15 minutes typically) from predefined geographic locations. If a check fails, you get an alert — often before any real user is affected.
Modern synthetic monitoring covers: uptime (is the site responding?), API endpoints (do our APIs return expected results?), browser transactions (can a user complete a purchase?), and SSL certificates (are they valid?).
Types of synthetic monitoring
Four check types that cover most needs
Uptime (HTTP/HTTPS) checks: GET a URL every N minutes; alert if it returns non-2xx or times out. The simplest form, best for public endpoints.
API monitoring: call an API with specific headers, body, and assertions on response code, body content, and latency. Essential for APIs that back mobile apps or partners.
Multi-step transactions: scripted workflows like "login → add to cart → checkout → confirm". Catches regressions in user flows that no single endpoint check would detect.
Browser-based monitoring: full browser (Puppeteer, Playwright) running a scripted scenario. Measures real page-load behavior including JS execution, Core Web Vitals, and third-party script performance.
Specialized checks: DNS resolution, SSL certificate expiry and chain validation, port/TCP checks, ICMP/ping. These are low-level checks for specific concerns.
Synthetic monitoring vs real user monitoring (RUM)
Two complementary approaches
RUM measures what your actual users experience: real browsers, real devices, real networks. Captures 100% of real traffic. Weakness: can't detect issues before users do.
Synthetic measures what a scripted probe experiences: controlled environment, consistent timing, predictable results. Strength: catches issues before users hit them, including during low-traffic periods (3 am).
Use synthetic for availability and consistency: did the login page work from Tokyo in the last 5 minutes? Use RUM for user experience: how slow is checkout for Android users on 3G?
Many teams run both. Synthetic is the leading indicator (fires first); RUM is the lagging indicator that confirms real-world impact.
Where to run checks: locations and frequency
Geography and schedule decisions
Run checks from locations that match your user base. If your users are in the US and Europe, run checks from US-East, US-West, London, and Frankfurt. Running only from one location hides regional outages.
Public endpoints: run from public cloud regions. Monitoring vendors offer 50–100 public probe locations.
Private/internal endpoints: run from a private location inside your network. Some vendors support self-hosted "private locations" that connect to the SaaS vendor via outbound tunnel.
Frequency: critical production endpoints every 1 minute; important endpoints every 5 minutes; low-priority endpoints every 15–30 minutes. Running every 10 seconds is overkill and creates noise.
Rotate locations between checks: if 3 locations check your API every minute, stagger them so you get a check from one of them every 20 seconds without hitting the API 3x more than needed.
Writing good synthetic checks
From "it exists" to "it works correctly"
Assert on content, not just status code. A page returning HTTP 200 with "Internal error" in the body is still broken. Assert that a specific expected string appears.
Measure the right latency. For a multi-step transaction, record and alert on each step's latency, not just the total. "The whole flow took 8s" is less actionable than "step 3 (payment call) took 6s".
Keep scripts idempotent. A browser transaction that creates test data must clean it up — or your production database will accumulate "test@monitoring.com" users.
Handle test accounts carefully. Monitoring from the public internet means your test credentials are on 50+ servers. Rotate frequently, scope narrowly, and never use production credentials.
Separate monitoring traffic from production analytics. Tag monitoring requests (e.g., User-Agent header) so they don't pollute conversion metrics.
Fail-fast on expected downstream: if a monitoring check depends on a third-party service, measure that dependency separately so you can distinguish "our site is down" from "third-party is down".
Synthetic monitoring tools
How to choose
Dedicated uptime tools: Pingdom, Uptime Robot, StatusCake, Better Uptime. Simple, cheap, good at HTTP checks. Fewer features for complex transactions.
Full APM/observability platforms with built-in synthetics: Datadog Synthetics, New Relic Synthetics, Dynatrace Synthetic, Atatus Uptime. Correlate synthetic failures with APM traces for faster diagnosis.
Open source: Checkly (cloud + self-hosted), Grafana Synthetic Monitoring (based on k6), Prometheus Blackbox Exporter. Lower cost; more operational overhead.
Specialized: Catchpoint for last-mile monitoring from ISP perspectives; ThousandEyes (Cisco) for network path monitoring.
Selection criteria: location coverage, scripting flexibility (API, multi-step, browser), alerting integrations, pricing per check per month, and correlation with your APM/log data.
Common patterns and anti-patterns
What good synthetic monitoring looks like
Pattern: cascading criticality. Core user journeys (login, checkout) get 1-minute browser checks. Core APIs get 1-minute API checks. Secondary endpoints get 5-minute checks. Internal admin get 15-minute.
Pattern: SSL expiry warning. Check SSL certificate expiry daily; alert at 30, 14, and 7 days remaining. Prevent the embarrassing weekend outage.
Pattern: canary deployment integration. Trigger synthetic checks against the canary environment after each deploy; promote only on green.
Anti-pattern: hundreds of overlapping uptime checks. Teams add checks over time and never prune. Review quarterly; delete checks for decommissioned services.
Anti-pattern: alerting from every location. If a check fails from 1 of 10 locations, that's probably a probe or ISP issue, not your service. Alert only when ≥ 2 locations agree.
Anti-pattern: synthetic as the only monitoring. Synthetic can't detect "the site is up but 20% of users are getting a 500". RUM and APM must complement it.
Key Takeaways
- Synthetic monitoring = scripted probes that simulate user interactions on a schedule.
- Four types: uptime, API, multi-step transaction, browser-based.
- Complements RUM: synthetic is the leading indicator, RUM is the user-experience ground truth.
- Run from multiple geographic locations; frequency matches criticality (1 min / 5 min / 15 min tiers).
- Assert on content and per-step latency; keep scripts idempotent; handle test credentials carefully.
- Alert only when multiple locations agree to avoid probe noise.
- Dedicated uptime tools for simple checks; integrated APM/observability platforms for correlation with traces and logs.