Skip to content

Domain controller monitoring ​

OpsMerge detects Windows domain controllers automatically and gives them monitoring built for the directory role, plus alerting that treats a DC going offline as the incident it is.

Automatic detection ​

A Windows agent reports whether it is a domain controller by reading its DomainRole (the reliable signal), rather than guessing from the OS name, which almost never says "domain controller". The first time an agent is detected as a DC, OpsMerge:

  • classifies it as a Domain Controller asset (promoting it even if it was previously mis-labelled a standard server), and
  • provisions a Domain Controller Health check on it, with no action needed from you.

Nothing to enable. A newly built or newly enrolled DC starts being monitored on its next check-in.

The Domain Controller Health check ​

Runs every 5 minutes on each DC and grades the directory's health in one place:

SignalSeverity if failing
Core services: NTDS, Netlogon, KDC, DNS runningError
AD replication to each partner succeedingError
SYSVOL and NETLOGON shares publishedError
AD database volume free space (the drive holding ntds.dit)Warning below 10%, Error below 5%
Clock synchronised (Kerberos breaks outside a 5-minute skew)Error
Support services: DFSR, W32Time runningWarning

The check reports one clear status and a message naming exactly what is wrong, for example "replication failing to 2 partner(s); AD database volume C: low (8% / 12GB free)". Signals it cannot measure (for instance replication on a box without the AD tools) are reported as unknown and never raise a false alert. Thresholds for the database volume are adjustable on the check like any other.

Timeouts ​

The probe runs as a single PowerShell pass with a 120-second budget. The replication query, which goes through Active Directory Web Services and can stall on a busy controller, runs inside its own 45-second job: if it does not answer in time the probe reports replication as not evaluated and still grades services, shares, disk and time. A probe that exceeds the whole budget is reported as "timed out after 120s" rather than "no output", so a slow controller is distinguishable from a broken one.

Alerting ​

A domain controller that goes offline always raises an availability alert, regardless of the per-device offline-alert setting. The offline-alert opt-out exists for devices that legitimately power off; a DC is never one of those. Critical servers get the same treatment.

The health check raises an alert when it fails, at the severity above, so a replication break, a stopped NTDS service, or a filling AD database volume pages you before it becomes an outage.

What it does not do ​

  • Non-Windows machines never carry this check.
  • It reads health, it does not change AD configuration.
  • It complements, rather than replaces, the general server monitoring checks (CPU, memory, disk, services) you would run on any server.

Lovingly Created in the UK
OpsMerge is a product of Brindleford Technologies Ltd, company number 16871436, registered in England and Wales.