A few days ago we wrote that companies should publish their own incident reports. Fair enough. Here's ours.
What happened
From the June migration of this site onto our own CMS until September 17, every one of the 100 posts in our feed rendered as an empty page. The feed itself looked fine — titles, images, excerpts, all there. Click any of them and you got the site header, a black screen, and nothing else. Old links from search engines redirected correctly… to a blank page. For three and a half months.
Why
Three small things lined up.
A naming convention was treated as a rule. When the core builds a post page it assembles it from three pieces — the theme's header, its post body, its footer — and it guessed their names from the theme's name. That guess was right for one theme and wrong for ours, which uses a different prefix. Another theme had hit the same bug earlier and worked around it by duplicating folders under the guessed names, instead of fixing the guess.
The renderer fails quietly, by design. If a page asks for a section type that doesn't exist, the core skips it rather than taking the page down. That's the right behaviour for a missing optional section. It's the wrong behaviour when all three sections are missing and the page is now empty.
The test that should have caught it couldn't fail. Our offline harness renders every page and post and reports PASS unless a template throws an error. Zero sections rendered, zero errors thrown: PASS, one hundred times.
Why nobody noticed
The feed index worked, so the site looked healthy at a glance. Nothing monitored whether a post had a body. And the honest answer: we don't read our own blog as often as we should.
What we changed
- The core now resolves the header, post and footer from the theme's actual section folders instead of guessing from its name. Any theme, any prefix.
- The harness fails a page that renders zero sections. A test that can't fail is decoration.
- The feed now reads the posts folder directly — it turned out a newly written post couldn't appear in the feed at all — and old links redirect in one hop instead of two.
- Next on the list: the ten-minute monitor will fetch a sample post and fail if the body is empty.
The part where we made it worse
While clearing caches after the fix, we deleted the cache directory itself instead of the files inside it. Nginx expected that directory to exist, so for about forty seconds every site on the server — not just ours — returned an error page. We recreated the folder and restarted the web server within the minute. Nobody wrote in. We're writing it here anyway, because that's the deal.
Lessons, in one line each
- A naming convention is not a contract. Look the thing up.
- "Skip silently" needs a floor: skipping everything is not a page.
- A green test suite that can't go red is a lie you tell yourself.
- Check what the visitor sees, not what the index says.
- Purge files, never folders.
If you're a client reading this and wondering whether your site had the same bug: no — it was specific to how this site's theme is named, and the fix went to every site the same day.