---
title: 12.01 Monitoring SimpleRisk
description: Monitor SimpleRisk via the healthcheck endpoint, web server logs, the debug log, the database (size, performance, replication if applicable), and the…
---

[Skip to content](https://support.simplerisk.com/kb/12-01-monitoring-simplerisk#main-content)

English

Show submenu for translations

[Customer portal](https://support.simplerisk.com/tickets?hsLang=en)

[![SimpleRisk logo of a man walking a tight rope](https://support.simplerisk.com/hs-fs/hubfs/simplerisk_logo_long_small-4.png?width=377&height=72&name=simplerisk_logo_long_small-4.png)](https://www.simplerisk.com/)

Open main navigation

Close main navigation

- English
  
  Show submenu for translations
- [Customer portal](https://support.simplerisk.com/tickets)
- [Contact us](https://www.simplerisk.com/about-us/contact-us)

[Contact us](https://www.simplerisk.com/about-us/contact-us)

 How can we help you?

- There are no suggestions because the search field is empty.

1. [SimpleRisk Knowledge Base](https://support.simplerisk.com/kb?hsLang=en)
2. [Administrator Guide](https://support.simplerisk.com/kb/administrator-guide?hsLang=en)
3. [12 Operations](https://support.simplerisk.com/kb/administrator-guide?hsLang=en#12-operations)

# 12.01 Monitoring SimpleRisk

## Monitor SimpleRisk via the healthcheck endpoint, web server logs, the debug log, the database (size, performance, replication if applicable), and the cron job runs. Forward logs to a SIEM, configure uptime monitoring on the healthcheck, alert on database growth and cron-job failures. Most issues surface in the logs before users report them.

## Why this matters

A monitored SimpleRisk install catches problems before users do. A SimpleRisk install nobody monitors fails silently in the night and produces "we're paying for this and it's broken" surprises in the morning. The infrastructure to monitor SimpleRisk is mostly off-the-shelf (uptime monitors, log aggregators, database monitoring tools), but it has to be configured. This article walks through what to monitor and how.

The honest scope to know up front: **SimpleRisk doesn't have a comprehensive built-in monitoring dashboard**. There's a healthcheck endpoint, the audit trail, and the debug log, but no "system health" page that summarizes everything. Operators stitch together the monitoring picture from external tools that consume SimpleRisk's signals.

## Before you start

Have these in hand:

- **Operational ownership of the SimpleRisk install** — who's on call, who responds to alerts, what's the escalation path.
- **A monitoring stack** — at minimum: an uptime monitor (Pingdom, UptimeRobot, Datadog Synthetics, or self-hosted Uptime Kuma); ideally also a log aggregator (Splunk, ELK, Datadog Logs, Grafana Loki) and an APM/metrics platform (Datadog, New Relic, Prometheus + Grafana).
- **Access to relevant infrastructure** — the SimpleRisk server (OS-level metrics), the database server (database metrics), the web server (access and error logs).

## What to monitor

### 1. Application reachability (uptime)

The healthcheck endpoint at `/healthcheck.php` returns a quick status response. Configure your uptime monitor to:

- **Hit `https://your-simplerisk.example.com/healthcheck.php`** every 1-5 minutes.
- **Expect a 200 response** with the expected body.
- **Alert on any non-200 response** or timeout.

If the healthcheck returns errors but the application looks fine to users, that's a configuration drift worth investigating; if the healthcheck succeeds but users report errors, the healthcheck might not cover the failing path.

### 2. Web server response time

Beyond uptime, response time matters. A page that loads in 30 seconds is "up" but unusable. Configure the uptime monitor (or a synthetic-monitoring tool) to:

- **Track response time** on the healthcheck and on a representative page (e.g., the login page).
- **Alert on response time exceeding a threshold** (e.g., 5 seconds for the healthcheck, 10 seconds for the login page).

### 3. Web server logs

Forward Apache or nginx access logs and error logs to your aggregator:

- **Access logs** — useful for traffic patterns, slow requests, error rates by endpoint.
- **Error logs** — PHP errors, web server errors. Alert on error rate spikes.

Common log patterns to alert on:

- High 5xx response rate (server errors).
- Sudden drop in successful response volume (the application is silently failing).
- Repeated requests for non-existent paths from a single source (scanning / probing).

### 4. The debug log

See [The Debug Log](https://support.simplerisk.com/kb/11-02-the-debug-log?hsLang=en). Monitor:

- **`error` and `critical` entries** — alert immediately. These indicate real problems.
- **`warning` entries** — review periodically; spikes may indicate emerging issues.
- **`notice` entries for failed cron runs** — alert on repeated failures of any cron job.

For installs writing to file or syslog, use the standard log-aggregation pipeline. For database destinations, periodic SQL queries can drive alerting.

### 5. Cron job execution

The cron jobs are SimpleRisk's background workers. Their failure produces visible application degradation (notifications stop sending, AI jobs queue indefinitely, workflows don't fire). Monitor:

- **Last successful run timestamp** for each cron job. Query `cron_history` (the table that records cron executions). Alert if any job hasn't run in N expected intervals.
- **Cron worker queue depth** — for the queue worker that processes background jobs, monitor pending job count. Sudden increases indicate the worker is falling behind.
- **Specific cron output files** — some installs write per-job logs. Tail and parse for failure patterns.

For installs running cron via system cron (not via the application's internal scheduler), the system cron itself may produce errors visible via `journalctl -u cron` or `mailx` to the operator account.

### 6. Database

Database health is application health. Monitor:

- **Connection pool utilization** — if SimpleRisk's connection pool is saturated, requests block.
- **Query performance** — slow queries log to MySQL's slow query log; ingest into the aggregator.
- **Replication lag** (if you have replication) — replicated reads from a lagged replica produce stale data.
- **Disk space** — running out of database disk is catastrophic. Alert at 80%, page at 90%.
- **Table size growth** — `audit_log` and `debug_log` grow continuously; monitor.

Database vendors (MySQL Enterprise Monitor, Percona Monitoring and Management) and APM tools have built-in database monitoring; configure for your install.

### 7. Disk space (application server)

The application server has logs, file uploads, temporary files. Monitor:

- **`/var/log/` and `simplerisk/logs/`** — log files grow without rotation.
- **`simplerisk/uploads/`** (if file uploads are stored locally) — user-uploaded documents accumulate.
- **`/tmp/`** — temporary files including activation backups. Should be cleaned up but sometimes aren't.
- **System root and database volumes** — generic disk-fill monitoring.

### 8. Memory and CPU

Standard server metrics:

- **CPU utilization** — sustained high CPU indicates load or runaway process.
- **Memory utilization** — Apache/nginx + PHP-FPM memory consumption; OOM-killer risk.
- **Swap usage** — sustained swap means insufficient memory.

These are operating system-level metrics; standard infrastructure monitoring covers them.

### 9. Application-specific metrics

Beyond infrastructure, track application-level metrics for the program:

- **Active user count** — how many users have authenticated in the last hour / day.
- **Risk submission rate** — risks created per day; sudden drops or spikes are worth investigating.
- **Job queue depth** — pending workflows, AI jobs, notification queue.
- **Authentication failure rate** — sustained high authentication failures may indicate brute-force.

These metrics typically come from SQL queries against the database run on a schedule and pushed to your metrics platform.

### 10. Backup verification

Backups that aren't tested aren't backups. Monitor:

- **Backup completion** — alert on missed backup runs.
- **Backup file size** — sudden change indicates corruption or scope change.
- **Periodic restore tests** — schedule actual restore drills (monthly or quarterly); alert on failures.

See [Database Backup and Restore](https://support.simplerisk.com/kb/02-05-database-backup-and-restore?hsLang=en).

## Alert thresholds and runbooks

For each alert type, define:

- **Threshold** — when does the alert fire?
- **Severity** — page (immediate response) vs ticket (next business day).
- **Runbook** — a documented procedure for diagnosing and resolving.

Without runbooks, alerts produce confused responders. Even a one-paragraph runbook ("when this fires, check X then Y, escalate to Z if not resolved in 30 minutes") materially improves incident response.

## Common operational signals

A handful of signal patterns recur:

- **Healthcheck failing**: SimpleRisk is down or degraded. Check web server, PHP-FPM, database connectivity.
- **Cron jobs not running**: workflows, notifications, AI all stop. Check cron daemon, system clock, application's cron worker process.
- **Database disk fill**: alert; truncate logs if appropriate; expand storage.
- **Login failures spiking**: brute-force attempt, credential leak, or systemic issue (LDAP outage, SSO problem).
- **Slow page loads**: database performance, web server tuning, opcache miss rate.
- **Notifications not sending**: SMTP connectivity, notification cron, queue depth.

## Common pitfalls

A handful of patterns recur with monitoring.

- **Configuring uptime monitoring without monitoring response time.** Slow but technically up is broken from the user's perspective.
- **Only monitoring what you know to monitor.** New failure modes appear after upgrades or feature additions. Periodically review what you're monitoring and what you're missing.
- **Alert fatigue from too-low thresholds.** Alerts that fire constantly get ignored. Tune thresholds to actual operational signals.
- **No runbooks for alerts.** A page at 3 AM with no runbook is a confused responder. Write runbooks for every page-level alert.
- **Not monitoring the monitoring system.** A dead Datadog agent doesn't alert that it's dead. Cross-monitor.
- **Treating the audit log as monitoring.** It captures changes, not state. Use the debug log + infrastructure metrics for monitoring.
- **Not testing alerts.** Configure an alert that's never been verified to actually fire — when it should fire in production, it doesn't. Test in non-production.
- **Storing all log data forever.** Storage cost compounds. Define retention; rotate old data to cheaper storage or delete.
- **Not monitoring backups.** Backups that fail silently lose you data when you need it.
- **Forgetting to monitor the database**. Database issues underlie most application issues. Monitor it explicitly.
- **Monitoring only via dashboards.** Dashboards require active viewing; alerts push to responders. Both have a place.

## Related

- [Performance Tuning](https://support.simplerisk.com/kb/12-02-performance-tuning?hsLang=en)
- [Scaling Considerations](https://support.simplerisk.com/kb/12-03-scaling-considerations?hsLang=en)
- [Troubleshooting Common Issues](https://support.simplerisk.com/kb/12-04-troubleshooting-common-issues?hsLang=en)
- [The Debug Log](https://support.simplerisk.com/kb/11-02-the-debug-log?hsLang=en)
- [The Audit Trail](https://support.simplerisk.com/kb/11-01-the-audit-trail?hsLang=en)
- [The Cron Jobs](https://support.simplerisk.com/kb/02-07-the-cron-jobs?hsLang=en)
- [Database Backup and Restore](https://support.simplerisk.com/kb/02-05-database-backup-and-restore?hsLang=en)
- [Log Rotation and Disk Management](https://support.simplerisk.com/kb/02-06-log-rotation-and-disk-management?hsLang=en)
- [Securing the Web Server](https://support.simplerisk.com/kb/09-05-securing-the-web-server?hsLang=en)

- [FAQs](https://support.simplerisk.com/kb/faqs?hsLang=en)
- [SimpleRisk Extras](https://support.simplerisk.com/kb/simplerisk-extras?hsLang=en#main-content)

    - [Vulnerability Management Extra](https://support.simplerisk.com/kb/simplerisk-extras?hsLang=en#vulnerability-management-extra)
    - [Team Separation Extra](https://support.simplerisk.com/kb/simplerisk-extras?hsLang=en#team-separation-extra)
    - [Import-Export Extra](https://support.simplerisk.com/kb/simplerisk-extras?hsLang=en#import-export-extra)
- [Administrator Guide](https://support.simplerisk.com/kb/administrator-guide?hsLang=en#main-content)

    - [00 About This Guide](https://support.simplerisk.com/kb/administrator-guide?hsLang=en#00-about-this-guide)
    - [01 Installation and Deployment](https://support.simplerisk.com/kb/administrator-guide?hsLang=en#01-installation-and-deployment)
    - [02 Upgrades and Maintenance](https://support.simplerisk.com/kb/administrator-guide?hsLang=en#02-upgrades-and-maintenance)
    - [03 Users and Permissions](https://support.simplerisk.com/kb/administrator-guide?hsLang=en#03-users-and-permissions)
    - [04 Authentication](https://support.simplerisk.com/kb/administrator-guide?hsLang=en#04-authentication)
    - [05 Customization](https://support.simplerisk.com/kb/administrator-guide?hsLang=en#05-customization)
    - [06 Configuring Risk and Compliance](https://support.simplerisk.com/kb/administrator-guide?hsLang=en#06-configuring-risk-and-compliance)
    - [07 Integrations](https://support.simplerisk.com/kb/administrator-guide?hsLang=en#07-integrations)
    - [08 The API](https://support.simplerisk.com/kb/administrator-guide?hsLang=en#08-the-api)
    - [09 Encryption and Data Security](https://support.simplerisk.com/kb/administrator-guide?hsLang=en#09-encryption-and-data-security)
    - [10 Workflows and Automation](https://support.simplerisk.com/kb/administrator-guide?hsLang=en#10-workflows-and-automation)
    - [11 Reporting and Auditing](https://support.simplerisk.com/kb/administrator-guide?hsLang=en#11-reporting-and-auditing)
    - [12 Operations](https://support.simplerisk.com/kb/administrator-guide?hsLang=en#12-operations)
    - [13 Reference](https://support.simplerisk.com/kb/administrator-guide?hsLang=en#13-reference)
- [User Guide](https://support.simplerisk.com/kb/user-guide?hsLang=en#main-content)

    - [00 Foundations of GRC](https://support.simplerisk.com/kb/user-guide?hsLang=en#00-foundations-of-grc)
    - [01 Risk Management](https://support.simplerisk.com/kb/user-guide?hsLang=en#01-risk-management)
    - [02 Compliance Management](https://support.simplerisk.com/kb/user-guide?hsLang=en#02-compliance-management)
    - [03 Governance](https://support.simplerisk.com/kb/user-guide?hsLang=en#03-governance)
    - [04 Asset and Data Inventory](https://support.simplerisk.com/kb/user-guide?hsLang=en#04-asset-and-data-inventory)
    - [05 Threat and Vulnerability Management](https://support.simplerisk.com/kb/user-guide?hsLang=en#05-threat-and-vulnerability-management)
    - [06 Assessments](https://support.simplerisk.com/kb/user-guide?hsLang=en#06-assessments)
    - [07 Incident Management](https://support.simplerisk.com/kb/user-guide?hsLang=en#07-incident-management)
    - [08 Audit and Reporting](https://support.simplerisk.com/kb/user-guide?hsLang=en#08-audit-and-reporting)
    - [09 Day-to-Day SimpleRisk](https://support.simplerisk.com/kb/user-guide?hsLang=en#09-day-to-day-simplerisk)
    - [10 Continuous Improvement](https://support.simplerisk.com/kb/user-guide?hsLang=en#10-continuous-improvement)
- [SimpleRisk User Guides](https://support.simplerisk.com/kb/simplerisk-user-guides?hsLang=en)
- [Troubleshooting](https://support.simplerisk.com/kb/troubleshooting?hsLang=en)
- [SimpleRisk Hosted](https://support.simplerisk.com/kb/simplerisk-hosted?hsLang=en)
- [How To videos](https://support.simplerisk.com/kb/how-to-videos?hsLang=en)

[![favicon-1](https://support.simplerisk.com/hs-fs/hubfs/favicon-1.png?width=35&height=35&name=favicon-1.png "favicon-1")](https://www.simplerisk.com)

<https://www.facebook.com/simplerisk/> <https://www.twitter.com/simpleriskfree/> <https://www.linkedin.com/company/simplerisk/>

Copyright © 2026, SimpleRisk, Inc.