Appearance
Domain controller monitoring
OpsMerge detects Windows domain controllers automatically and gives them monitoring built for the directory role, plus alerting that treats a DC going offline as the incident it is.
Automatic detection
A Windows agent reports whether it is a domain controller by reading its DomainRole (the reliable signal), rather than guessing from the OS name, which almost never says "domain controller". The first time an agent is detected as a DC, OpsMerge:
- classifies it as a Domain Controller asset (promoting it even if it was previously mis-labelled a standard server), and
- provisions a Domain Controller Health check on it, with no action needed from you.
Nothing to enable. A newly built or newly enrolled DC starts being monitored on its next check-in.
The Domain Controller Health check
Runs every 5 minutes on each DC and grades the directory's health in one place:
| Signal | Severity if failing |
|---|---|
| Core services: NTDS, Netlogon, KDC, DNS running | Error |
| AD replication to each partner succeeding | Error |
| SYSVOL and NETLOGON shares published | Error |
| AD database volume free space (the drive holding ntds.dit) | Warning below 10%, Error below 5% |
| Clock synchronised (Kerberos breaks outside a 5-minute skew) | Error |
| Support services: DFSR, W32Time running | Warning |
The check reports one clear status and a message naming exactly what is wrong, for example "replication failing to 2 partner(s); AD database volume C: low (8% / 12GB free)". Signals it cannot measure (for instance replication on a box without the AD tools) are reported as unknown and never raise a false alert. Thresholds for the database volume are adjustable on the check like any other.
Timeouts
The probe runs as a single PowerShell pass with a 120-second budget. The replication query, which goes through Active Directory Web Services and can stall on a busy controller, runs inside its own 45-second job: if it does not answer in time the probe reports replication as not evaluated and still grades services, shares, disk and time. A probe that exceeds the whole budget is reported as "timed out after 120s" rather than "no output", so a slow controller is distinguishable from a broken one.
Alerting
A domain controller that goes offline always raises an availability alert, regardless of the per-device offline-alert setting. The offline-alert opt-out exists for devices that legitimately power off; a DC is never one of those. Critical servers get the same treatment.
The health check raises an alert when it fails, at the severity above, so a replication break, a stopped NTDS service, or a filling AD database volume pages you before it becomes an outage.
What it does not do
- Non-Windows machines never carry this check.
- It reads health, it does not change AD configuration.
- It complements, rather than replaces, the general server monitoring checks (CPU, memory, disk, services) you would run on any server.