In any environment that exposes endpoints, it is important to make sure they are both secured and monitored. There is plenty of documentation on performance monitoring, but very little on monitoring the security of endpoints, and specifically TLS certificates on non-HTTP endpoints. This post shows how to use the Prometheus Blackbox Exporter to do exactly that.
For example, when we deploy Apache Cassandra, it is typical to secure the database endpoints with certificates so that data is encrypted in flight and, if required, client certificates are validated. Knowing when these certificates expire, and being able to monitor and alert on it, is critical. The same approach applies to any application that exposes a TLS endpoint, such as LDAP, Kafka or ELK.
For simplicity, we will use a single Blackbox probe on the same VM as a single Prometheus instance, monitoring the certificate on one Apache Cassandra node.
Check the certificate manually
First, check the connection from the Prometheus host to our node, cas01.dev.db.myexample.io, on port 9142 (Cassandra's TLS port), using OpenSSL:
root@prom.blog.myexample.io:~# echo -n | openssl s_client -connect cas01.dev.db.myexample.io:9142 2> /dev/null | openssl x509 -noout -text
Certificate:
Data:
Version: 1 (0x0)
Serial Number: 13315684638806572674 (0xb8cad0a12530fe82)
Signature Algorithm: sha256WithRSAEncryption
Issuer: C=US, O=myexample.io, OU=DevCluster, CN=rootCA
Validity
Not Before: Jan 22 13:01:15 2019 GMT
Not After : Jan 21 13:01:15 2021 GMT
Subject: C=US, O=myexample.io, OU=DevCluster, CN=cas01.dev.db.myexample.io
[...]
echo -n makes OpenSSL return to the prompt immediately, openssl s_client connects to the endpoint and reads the certificate, and openssl x509 displays it. Keep this command in your toolbox: it is the same one you need when Prometheus reports a failing probe and you want to confirm the actual state of the endpoint.
Configure the Prometheus Blackbox Exporter
Configure a TCP module on the Blackbox Exporter that opens a TLS connection, in blackbox.yml:
modules:
tcp_cert:
prober: tcp
timeout: 5s
tcp:
tls: true
tls_config:
insecure_skip_verify: true
You can test the probe directly:
root@prom.blog.myexample.io:~# curl 'http://prom.blog.myexample.io:9115/probe?target=cas01.dev.db.myexample.io:9142&module=tcp_cert&debug=true'
Against a real Cassandra node, this returns the metrics we are looking for:
ts=2020-08-01T09:39:41.498464359Z caller=main.go:304 module=tcp_cert target=cas01.dev.db.myexample.io:9142 level=info msg="Beginning probe" probe=tcp timeout_seconds=5
ts=2020-08-01T09:39:41.498568598Z caller=tcp.go:41 module=tcp_cert target=cas01.dev.db.myexample.io:9142 level=info msg="Resolving target address" ip_protocol=ip6
ts=2020-08-01T09:39:41.503386291Z caller=tcp.go:41 module=tcp_cert target=cas01.dev.db.myexample.io:9142 level=info msg="Resolved target address" ip=10.0.4.20
ts=2020-08-01T09:39:41.503413503Z caller=tcp.go:111 module=tcp_cert target=cas01.dev.db.myexample.io:9142 level=info msg="Dialing TCP with TLS"
ts=2020-08-01T09:39:41.524602341Z caller=main.go:119 module=tcp_cert target=cas01.dev.db.myexample.io:9142 level=info msg="Successfully dialed"
ts=2020-08-01T09:39:41.524669931Z caller=main.go:304 module=tcp_cert target=cas01.dev.db.myexample.io:9142 level=info msg="Probe succeeded" duration_seconds=0.026147547
Metrics that would have been returned:
# HELP probe_dns_lookup_time_seconds Returns the time taken for probe dns lookup in seconds
# TYPE probe_dns_lookup_time_seconds gauge
probe_dns_lookup_time_seconds 0.004831258
# HELP probe_duration_seconds Returns how long the probe took to complete in seconds
# TYPE probe_duration_seconds gauge
probe_duration_seconds 0.026147547
# HELP probe_failed_due_to_regex Indicates if probe failed due to regex
# TYPE probe_failed_due_to_regex gauge
probe_failed_due_to_regex 0
# HELP probe_ip_protocol Specifies whether probe ip protocol is IP4 or IP6
# TYPE probe_ip_protocol gauge
probe_ip_protocol 4
# HELP probe_ssl_earliest_cert_expiry Returns earliest SSL cert expiry date
# TYPE probe_ssl_earliest_cert_expiry gauge
probe_ssl_earliest_cert_expiry 1.611234074e+09
# HELP probe_success Displays whether or not the probe was a success
# TYPE probe_success gauge
probe_success 1
# HELP probe_tls_version_info Returns the TLS version used, or NaN when unknown
probe_tls_version_info{version="TLS 1.2"} 1
Module configuration:
prober: tcp
timeout: 5s
http:
ip_protocol_fallback: true
tcp:
ip_protocol_fallback: true
tls: true
tls_config:
insecure_skip_verify: true
icmp:
ip_protocol_fallback: true
dns:
ip_protocol_fallback: true
The probe works: it connects to the Cassandra node and exports the metrics we need. The most useful are:
probe_success: confirms the endpoint is reachable;probe_ssl_earliest_cert_expiry: the expiry time of the certificate, as a Unix timestamp;probe_duration_seconds: how long the probe took, useful for checking the responsiveness of the node.
The Blackbox Exporter also exposes probe_ssl_last_chain_info, a metric that carries the leaf certificate's fingerprint, issuer, subject and SAN list as labels. It is handy for spotting an unexpected certificate swap (wrong CA, wrong CN) even when the expiry date itself still looks fine.
Why insecure_skip_verify is enabled
The probe above sets insecure_skip_verify: true. That is deliberate: without it, the probe only succeeds against a certificate that is currently valid and correctly chained to a CA the exporter trusts. With it, you still get a reading, and an accurate expiry timestamp, from a certificate that has already expired or was issued by a CA you have not imported into the exporter's trust store.
If you also want to know whether the certificate is actually trusted, add a second, verifying module alongside the first:
modules:
tcp_cert_verified:
prober: tcp
timeout: 5s
tcp:
tls: true
tls_config:
insecure_skip_verify: false
ca_file: /etc/blackbox_exporter/rootCA.crt
Verification also checks the hostname, and the Blackbox Exporter (like any modern Go program) only matches it against the certificate's Subject Alternative Name (SAN) list. A certificate that names the host only in its CN, such as the version 1 certificate in the example above, will fail this probe even if it is otherwise valid.
Both modules are then scraped by Prometheus, as shown below.
Configure Prometheus
In prometheus.yml, add one scrape job per module:
scrape_configs:
- job_name: blackbox_cas
params:
module:
- tcp_cert
metrics_path: /probe
static_configs:
- targets:
- cas01.dev.db.myexample.io:9142
relabel_configs:
- source_labels: [__address__]
target_label: __param_target
- source_labels: [__param_target]
target_label: instance
- target_label: __address__
replacement: 127.0.0.1:9115
- job_name: blackbox_cas_verified
params:
module:
- tcp_cert_verified
metrics_path: /probe
static_configs:
- targets:
- cas01.dev.db.myexample.io:9142
relabel_configs:
- source_labels: [__address__]
target_label: __param_target
- source_labels: [__param_target]
target_label: instance
- target_label: __address__
replacement: 127.0.0.1:9115
cas01.dev.db.myexample.io:9142 is our node's hostname and port; in this test environment it is the only target. 127.0.0.1:9115 is the address of the Blackbox Exporter, which here runs on the same host as Prometheus on its standard port.
Configure the alert rules
With the metrics in place, define the alerts. The expiry rules read only the non-verifying job, so each certificate is reported once, and the two failure alerts separate an unreachable endpoint from an untrusted certificate:
groups:
- name: tls-cert-expiry
rules:
- alert: BlackboxSSLCertificateWillExpireVerySoon
expr: 0 <= (last_over_time(probe_ssl_earliest_cert_expiry{job="blackbox_cas"}[10m]) - time()) / 86400 < 3
labels:
severity: critical
annotations:
summary: "SSL certificate for {{ $labels.instance }} expires in less than 3 days"
- alert: BlackboxSSLCertificateWillExpireSoon
expr: 3 <= (last_over_time(probe_ssl_earliest_cert_expiry{job="blackbox_cas"}[10m]) - time()) / 86400 < 20
labels:
severity: warning
annotations:
summary: "SSL certificate for {{ $labels.instance }} expires in less than 20 days"
- alert: BlackboxSSLCertificateExpired
expr: (last_over_time(probe_ssl_earliest_cert_expiry{job="blackbox_cas"}[10m]) - time()) < 0
labels:
severity: critical
annotations:
summary: "SSL certificate for {{ $labels.instance }} has already expired"
- alert: BlackboxProbeFailed
expr: probe_success{job="blackbox_cas"} == 0
for: 2m
labels:
severity: critical
annotations:
summary: "Blackbox probe failed for {{ $labels.instance }}: endpoint unreachable or TLS handshake failed"
- alert: BlackboxSSLCertificateUntrusted
expr: probe_success{job="blackbox_cas_verified"} == 0 and on(instance) probe_success{job="blackbox_cas"} == 1
for: 2m
labels:
severity: warning
annotations:
summary: "SSL certificate for {{ $labels.instance }} is reachable but fails verification (expired, untrusted CA or hostname mismatch)"
Configure Grafana
To visualise the metrics, you can use community dashboards from the Grafana website, such as dashboard 7587 or dashboard 11529.
Closing
Certificate monitoring is not sufficient on its own for high-availability or production environments, but it should be part of every monitoring system. No one likes being woken in the middle of the night because production is down due to an expired certificate. Make sure your monitoring displays, and alerts on, certificate expiry.



