Proxy 05 - Certificate renewal
Proxy Solution · Previous: Squid with SSL interception · Next: SELinux and client trust
A proxy that bumps TLS signs a certificate for every server name its clients visit, and it signs them with a certificate of its own that every client has to trust. That certificate was self-signed and, in production, valid for one year, so when it ran out it had to be renewed by hand on each proxy and then carried to every server that trusts it; the notes hold one renewal per host. This Article is about that certificate: why it is self-signed, where it lives on the lab host and on the two production hosts, the renewal sequence as it was run, and the part the notes do not hold, the deployment to the servers. The command sets are Config documents: the lab CA certificate commands, certificate renewal on dc1-a-vcprx001 and certificate renewal on dc2-a-vcprx001.
Why a self-signed certificate
The notes do not argue it; they start with openssl genrsa. The argument, as I understand SSL bumping, is this: with generate-host-certificates=on Squid's ssl_crtd helper creates a certificate for pypi.python.org, or whatever server name the client asked for, on the fly, and signs it with the key and certificate the cert= and key= options of the port name. A public CA will not issue a certificate that may sign other certificates to a proxy, and no public CA is in a position to vouch for a certificate that claims to be pypi.python.org and is not. So the signing certificate has to be a private one, and the clients have to be told to trust it. That is a decision with two halves: the proxy gets a self-signed certificate, and every client that goes through the proxy gets that certificate in its trust store. The second half is what makes the renewal more than a local change on the proxy.
The certificate was made with openssl req and openssl x509 -req -signkey, with no extensions file. As I read openssl x509, that gives an X.509 version 1 certificate without basicConstraints, so nothing in it says "CA". Squid signed with it regardless. Whether the clients that received the system bundle accepted the chain the notes imply and do not show: the pip test of the next Article records the REQUESTS_CA_BUNDLE export as the solution and no second run after it. The notes do not go into it, and I did not check whether a stricter client would have minded.
Where the files are
The lab put the key under /etc/pki/tls/private/ and the certificate under /etc/pki/tls/certs/, the RHEL default places, and gave the certificate 3652 days, ten years and two days. Production put everything under /etc/ssl/squid/ with a suffix per datacenter, _1 and _2, and 365 days. The two production hosts do not spell the suffix the same way, and the notes do not say why; the number matches the datacenter, which is my reading of it.
| Host | Key | Certificate | CSR | Validity |
|---|---|---|---|---|
proxy.lab.example.net | /etc/pki/tls/private/proxy.lab.example.net.key | /etc/pki/tls/certs/proxy.lab.example.net.crt | /etc/pki/tls/private/proxy.lab.example.net.csr | 3652 days |
dc1-a-vcprx001 | /etc/ssl/squid/dc1-a-vcprx001.adm.example.net.private_1 | /etc/ssl/squid/dc1-a-vcprx001.adm.example.net.crt_1 | /etc/ssl/squid/dc1-a-vcprx001.adm.example.net.csr_1 | 365 days |
dc2-a-vcprx001 | /etc/ssl/squid/dc2-a-vcprx001_2.key | /etc/ssl/squid/dc2-a-vcprx001_2.crt | /etc/ssl/squid/dc2-a-vcprx001_2.csr | 365 days |
The file names come from the two port lines that grep ssl-bump found in each production squid.conf, an http_port 3128 ssl-bump … and an https_port 3128 … ssl-bump intercept … on the same port. Both are printed without a comment sign; the notes hold nothing else of the production squid.conf. The lines themselves, and the question of how Squid ran with both, are in Squid with SSL interception.
When they expired
The renewal starts with a look at the certificate's end date. On dc1-a-vcprx001, as root:
$ openssl x509 -in /etc/ssl/squid/dc1-a-vcprx001.adm.example.net.crt_1 -text -noout | grep "Not After"
output 1 line
Not After : Apr 9 09:00:00 2019 GMT
| Host | Not After |
|---|---|
dc1-a-vcprx001 | Apr 9 09:00:00 2019 GMT |
dc2-a-vcprx001 | Apr 27 13:13:55 2019 GMT |
With 365-day certificates, those two dates put the issue of the certificates then in place in April 2018, which fits the rollout of the proxies in 2018; that is my arithmetic, the notes carry no date in their text. The datacenter 1 note is filed under 2019, the datacenter 2 note under 2018; both carry the step title "Create new CSR (if certificate is expired)" and read as the procedure for their host, and the datacenter 2 note goes straight on into the SELinux problem of the next Article.
The renewal sequence
The sequence is the same on both hosts and the notes write every step with its own title. First the port lines are grepped to see which files are in use. Then the expiry date is checked. Then a new signing request is made from the existing key, so the key pair stays and only the certificate changes; openssl req asks the distinguished-name questions again and they are answered again, Common Name the host's FQDN; the original answers are not in the notes. Then the request is self-signed with the same key for 365 days, straight over the old certificate file. Then Squid is stopped, the certificate cache /var/lib/ssl_db is deleted, created again with ssl_crtd -c -s, its files given to the squid user, and Squid started. The middle of it, on dc2-a-vcprx001 as root:
$ openssl req -new -key /etc/ssl/squid/dc2-a-vcprx001_2.key -out /etc/ssl/squid/dc2-a-vcprx001_2.csr $ openssl x509 -req -days 365 -in /etc/ssl/squid/dc2-a-vcprx001_2.csr -signkey /etc/ssl/squid/dc2-a-vcprx001_2.key -out /etc/ssl/squid/dc2-a-vcprx001_2.crt $ systemctl stop squid $ rm -rf /var/lib/ssl_db $ /usr/lib64/squid/ssl_crtd -c -s /var/lib/ssl_db $ chown -R squid:squid /var/lib/ssl_db/* $ systemctl start squid
Why the cache goes: the notes only say "Remove Squid SSL cache". My inference is that /var/lib/ssl_db holds the host certificates that ssl_crtd generated and signed with the old certificate, and Squid would keep serving a cached pypi.python.org certificate chained to the old signer for as long as it stayed in the database. A client that has only the new certificate in its trust store, or whose old copy has passed its Not After, would then refuse exactly the servers it visits most. Deleting the database makes every host certificate be generated again under the new signer. As I read it, it is also the only step of the sequence that needs Squid stopped, which is why the stop and start bracket it and not the signing. The chown glob gives the squid user the files inside the directory and, as I read it, leaves the directory itself to root, as the lab did; the notes do not remark on it and end with systemctl start squid, with no output under it.
sequenceDiagram participant A as admin participant P as proxy host participant S as squid participant C as ssl_db cache participant D as colleague's deployment participant V as servers A->>P: grep ssl-bump, which files are in use A->>P: openssl x509 -noout, Not After A->>P: openssl req -new -key, new CSR A->>P: openssl x509 -req -days 365, new crt A->>S: systemctl stop squid A->>C: rm -rf /var/lib/ssl_db A->>C: ssl_crtd -c -s, chown squid A->>S: systemctl start squid A->>D: the new crt file D->>V: deployment, not in the notes
The hand-over
Above the signing step both notes carry the same NOTE: the new certificate file needs to be sent to the colleague who runs the certificate deployment, "Automatic Deployment of certificates to servers". That is all the notes hold of the second half of the decision from the first section. How the file travels, which Ansible role or other tool puts it into /etc/pki/ca-trust/source/anchors/ on each server and runs update-ca-trust, whether the servers get it before or after the proxy switches, and what happens to a server that is missed: none of it is in my notes, because none of it was my work. What the lab did on its own host, cp into the anchors and update-ca-trust, is what I would expect that deployment to do on every server, but that is an expectation, not a record.
The datacenter 1 NOTE names the file /etc/ssl/squid/dc1-a-vcprx001.adm.example.net.crt_1t. The signing command writes ….crt_1; the trailing t is a typo in the note, kept in the Config document and pointed out there, because a colleague who copies the name from the note would look for a file that does not exist.
What is missing
Beyond the deployment: the first creation of the production keys and certificates is not in the notes, only the renewal from an existing key. The CSR answers are shown with placeholders here. There is no check after systemctl start squid, no openssl x509 -noout on the new file, no client test after the renewal, and no reminder or monitoring of the Not After date; the Sensu hosts reach the proxies for SNMP, but what they check is not in the notes either. Why the production ports carry version=1 and options=NO_SSLv2,NO_SSLv3,SINGLE_DH_USE while the lab's did not is not written down.
Today
Checked against OpenSSL 3.5 and Squid 7.7: the two openssl commands still run (-signkey is an alias of -key since OpenSSL 3.0, and since 3.2 x509 emits version 3 certificates with key identifiers), but the certificate they make is the wrong kind for a current Squid, which documents that with generate-host-certificates=on "the first tls-cert= option must be a CA certificate capable of signing the automatically generated certificates"; the x509 -req path without extensions yields no basicConstraints, while the Squid wiki's one-step openssl req … -x509 -extensions v3_ca recipe, or req -x509 with the stock openssl.cnf, yields a CA certificate. My inference about the cache is confirmed by the wiki: "whenever you change the signing CA be sure to erase and re-initialize the certificate database. It contains signed certificates and clients may experience connectivity problems when the signing CA no longer matches your configured CA". Squid also documents that the generated host certificates live as long as the CA certificate, so a 365-day signing certificate means 365-day host certificates. The helper was renamed in Squid 4 to security_file_certgen, its default database moved to /var/spool/squid/ssl_db on RHEL builds, and -s now requires -M, so the re-creation reads security_file_certgen -c -s DIR -M 4MB. The platform is gone too: Squid 3.5 ended with 3.5.27 in August 2018, and RHEL 7 left Maintenance Support on 2024-06-30.
What I would do differently
Make the signing certificate a real CA certificate with the one-step req -x509 recipe, which a current Squid insists on anyway. Name the files the same way on both hosts, since the renewal is the same sequence and the one difference between the two notes is which name to type. And verify the new file's Not After and the service with one more command each before handing the certificate over; the notes stop at systemctl start squid, and the one typo they contain is in the hand-over line.