Sunday, October 11, 2026

Apache/PHP-FPM: all fun and games until someone uses Proxy

Recently at work, files started being owned by the wrong user.  Ultimately, I would find that the culprit was a buffering change unexpectedly interacting with SetHandler directives in Apache.

Some experimentation had demonstrated that PHP’s flushing functions were not working to generate output immediately, and the slop machine suggested adding a simple little config block:

<Proxy "fcgi://localhost">
    ProxySet flushpackets=on
</Proxy>

Once this is present, all SetHandler "...|fcgi://localhost" directives become “shared workers”, and Apache assumes that all the sockets defined in the SetHandler are alternate paths to the same destination. Requests can suddenly be routed to the wrong worker pool.  To fix it, our secondary SetHandler needed to carry a different ID than “localhost”, as in:

SetHandler "unix:proxy:/run/php/sns.sock|fcgi://deploy"

With different IDs, the destinations are treated as distinct once again.

When the slop machine gave me the <Proxy> block, it didn’t know that production had a separate FPM pool, also configured using SetHandler, with the same worker ID of “localhost”.  For my part, I didn’t know it was a worker ID.  It looked to me like it’s part of the thing that makes it work.

# system default config
SetHandler "unix:proxy:/run/php/php85-fpm.sock|fcgi://localhost"
# SNS deployment config (wrong)
SetHandler "unix:proxy:/run/php/sns.sock|fcgi://localhost"

For years, both of these pools have been running side-by-side, treated as unique destinations by Apache.  There was no indication that there was a conflicting/duplicate worker ID here.  The choice of localhost instead of pool-www or something in the default configuration led me to copy it for the secondary pool.  After all, PHP-FPM was running on localhost.

However, adding the <Proxy> block caused spooky action at a distance.  Apache no longer set up two separate pools.  Instead, it assumed a single, shared worker that just happened to be reachable through two different sockets in the filesystem.  Requests that had long been routed separately to the sockets as intended would now be arriving on whichever socket Apache happened to try first!

This immediately exploded for us because the two FPM pools run as different system users.  Standard websites do not have ownership over their document roots; they cannot rewrite their own files.  The SNS deployment system runs on a separate pool as a different user, which can alter the files.  That is its job.

In normal operation, the regular website maintains a cache that is private to the exact running user.  If the SNS pool picks up the request and generates files, they’ll be owned by the SNS/deployment user.  If Apache decides to use the other socket later, the cache is unexpectedly inaccessible.  Logging into the server can show the unexpected ownership, but doesn’t explain why it happened.

The fix was to give the SNS pool its own distinct worker ID, as demonstrated at the beginning of the post.  Once it was changed to fcgi://deploy, the configuration works as intended again.  We have one Apache worker (id localhost) dispatching normal websites on the www pool with its default socket, and another distinct worker (id deploy) dispatching to the sns pool.  Only the www pool is latency-sensitive, so it retains the flushpackets=on directive as above, and the other pool doesn’t get any special Apache configuration.

No comments: