SYS:susipere.de./en/journal/nginx-basic-auth-401-dead-upstream/

A dead service behind nginx basic auth still returns 401

An endpoint on our own server answered HTTP 401 for months with nothing running behind it. Why auth in front of a proxy hides a dead upstream.

We ran a read-only inventory of our production server on 26 September 2026 and found a public HTTPS endpoint that had been answering requests since April with nothing behind it at all.

It was not returning 502. It was returning 401.

What we found

The endpoint is a browser-proxy route on one of our own hostnames. nginx terminates TLS on a high port, requires HTTP basic auth, and proxies to a local port:

server {
    listen 18443 ssl;
    server_name  ...;

    auth_basic            "...";
    auth_basic_user_file  /etc/nginx/.htpasswd-...;

    location / {
        proxy_pass http://127.0.0.1:18789;
    }
}

Nothing has been listening on 127.0.0.1:18789 for months. There is no process and no systemd unit — either the service was never installed as a managed unit on this box, or it was removed and the route stayed behind.

From outside, the endpoint looks fine:

$ curl -s -o /dev/null -w '%{http_code} %{time_total}\n' https://.../
401 0.019678

Fifteen to twenty milliseconds, and a 401. That is a perfectly healthy-looking response from a service that does not exist.

On the host, the truth is one command away — if you know to look for it:

$ ss -ltnp | grep 18789
(no output)

$ curl http://127.0.0.1:18789/
curl: (7) Failed to connect to 127.0.0.1 port 18789: Connection refused

Why a 401 comes back instead of a 502

nginx handles a request in ordered phases. auth_basic is evaluated in the access phase. proxy_pass runs in the content phase, which is later. When the access phase rejects a request, the content phase never runs — so nginx never opens a connection to the upstream, and never discovers that the upstream is gone.

The tell is in the error log. A genuinely broken proxy writes a line like connect() failed (111: Connection refused) while connecting to upstream. Ours wrote nothing at all, because no connection was ever attempted.

Reproducing it in about a minute

This needs no application and no second machine. Write the config:

daemon off;
worker_processes 1;
events { worker_connections 16; }
http {
  access_log off;
  server {
    listen 127.0.0.1:18999;

    auth_basic            "probe";
    auth_basic_user_file  /tmp/nginx-401-repro/htpasswd;

    location / {
      proxy_pass http://127.0.0.1:18789;   # nothing is listening here
    }
  }
}

Create a password file and start a throwaway nginx in the foreground:

mkdir -p /tmp/nginx-401-repro
printf 'probe:%s\n' "$(openssl passwd -apr1 probe)" > /tmp/nginx-401-repro/htpasswd
chmod 644 /tmp/nginx-401-repro/htpasswd
nginx -p /tmp/nginx-401-repro -c /tmp/nginx-401-repro/nginx.conf \
      -e /tmp/nginx-401-repro/error.log

Then the two requests that tell the whole story:

# no credentials: nginx answers from the auth layer
$ curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1:18999/
401
$ grep -c upstream /tmp/nginx-401-repro/error.log
0

# valid credentials: the proxy is actually attempted
$ curl -s -o /dev/null -w '%{http_code}\n' -u probe:probe http://127.0.0.1:18999/
502
$ grep -c upstream /tmp/nginx-401-repro/error.log
1

Same config, same dead upstream, same instant. The status code depends entirely on whether the request got past the auth layer.

The actual lesson

Any check that sits in front of an authentication layer is testing the authentication layer. Ours reported a dead service as reachable for months, because "it responded and it wasn't a 5xx" was the entire test.

This generalises well past basic auth. The same false negative shows up with:

  • an IP allowlist returning 403
  • a CDN or WAF challenge page returning 403 or 503
  • a maintenance page served by the proxy itself
  • a /healthz answered by return 200 in the proxy config instead of being proxied to the application

In every one of those cases the proxy answers on the application's behalf, and the application's actual state is invisible.

Three fixes, in the order we would apply them:

  1. Assert on the expected response, not the absence of failure. Require a specific status code and a specific piece of the body. "Not a 5xx" is not a health check.
  2. Probe the service where it actually lives. From the host, curl http://127.0.0.1:18789/ answers the question the public URL cannot.
  3. Give the check a path through the auth layer. A location = /healthz with auth_basic off; that proxies to the application will return a real 502 when the application is down. Keep that endpoint's response to liveness only — it is unauthenticated by design.

Where this one stands

We have not fixed it yet, and that is the honest answer. The route is out of our "working" column, and the service is now an explicit delete-or-fix decision rather than an assumed-healthy dependency. A dead upstream is not itself exploitable, so this is a monitoring failure rather than an incident — which is exactly why it survived so long.

We do not know when the upstream died. We know the route config was written on 7 April 2026 and that nothing was behind it on 26 September 2026. Nothing alerted in the 172 days between, because nothing was ever asking the right question.

Reference: ngx_http_auth_basic_module and the phase ordering in the nginx development guide. Measurements taken on nginx 1.26.0 (Ubuntu).

cd /en/journal