software engineering
You can do everything right and still have a slow website
Load testing, misplaced confidence, and the questions you only think to ask after everything else is working.
Back at my old company we worked with one of Finland’s largest online adult-toy retailers. The relationship was pretty straightforward: we would put the site under increasing amounts of traffic, find where it started to fall apart, hand the findings to their IT team, and they would fix what we had exposed.
They had good reasons to care. People hate waiting when trying to buy something. Slow pages mean worse conversion rates and more people leaving. Google cares about it too. PageSpeed Insights and Core Web Vitals give you a pretty good idea of the sort of performance they measure, and page experience can affect how well a site performs in search. So for the retailer faster pages meant a better customer experience, better search visibility, and fewer abandoned purchases.
One of the interesting things you learn when working with a large adult-toy retailer is when people buy their toys. And according to our client's traffic data, it's usually in the evening. Traffic started climbing noticeably after around 20:00. Sunday through Tuesday were especially strong. During seasonal peaks, overall traffic was higher. Christmas, for example, was busy but relatively smooth, more customers but no huge spikes. People have weeks to do their Christmas shopping.
Black Friday was different. Black Friday could mean roughly ten times normal traffic. And one of their earlier Black Fridays had become something of a cautionary tale. They had prepared for the campaign and thought the infrastructure was ready. The promotion went live at around 21:00. Then the servers started crashing. The person who was supposed to be on call had his phone on silent and was, apparently, in the sauna. The store was down for roughly half an hour while a heavily promoted campaign was already live. Eventually they got hold of him, restarted the systems, and brought the store back.
The embarrassing part was not that something failed. Systems fail. The embarrassing part was that everyone had believed the platform could handle the traffic. That belief had simply never been tested against reality.
This is where we came in. Our goal as a company was to be the Black Friday before the Black Friday.
When we started working together, the webshop was still running on fairly traditional virtual servers. Under normal traffic everything looked fine, but once we got to around 300 concurrent users in the load test, it started choking.
The first problems were fairly mundane. Apache configuration, Tomcat connection limits, database connection limits. Things that had probably been sitting there quite happily for years because normal traffic had never given them a reason to become a problem. Their IT team changed the configuration and we tested again. The site could handle roughly twice as much traffic. We kept pushing it for a while after that, but eventually we were getting diminishing returns. There were still things we could tune, but we had more or less squeezed what we could out of the existing architecture without making bigger changes.
So they went away to think about it.
Some weeks later, their head of IT joined one of our meetings with a newly appeared, familiar blue helm sticker on his laptop. I had a fairly good idea where this was going. The platform was moving to Kubernetes.
It was much better. Even the preliminary tests were a pretty dramatic improvement. Where the old setup had started running into trouble somewhere around 600 concurrent users after all the tuning, the Kubernetes version got us into the thousands. At around 3,000 concurrent users we were still finding new limits rather than simply watching the whole thing collapse.
And this wasn't because they had solved the problem by throwing an absurd amount of hardware at it. The underlying servers were still roughly in the same class. Most of the improvement came from changing how the application was deployed and how the available capacity was being used.
Of course, Kubernetes also gave us new things to break. One of them was scaling itself. The cluster could add capacity when traffic increased, but adding capacity takes time. If we increased the load quickly enough, traffic could arrive faster than Kubernetes could bring new instances online.
For one of our last big tests we noticed something weird, images started falling behind and eventually some of them weren't arriving at all. That didn't look like the failures we'd been chasing up to that point. At this point there were not many sensible places left to look inside the application. So we started looking further down the stack, and after all of the options we could think of were exhausted, their head of IT called the datacenter and asked them to check the network side.
Turns out the connection was running at 100 Mbps. It was an issue with the wire. We had optimized our way all the way down to the wire.
Thinking back, there's an interesting contradiction in all of this. One of our starting assumptions was that old-school load testing wasn't enough, because it doesn't reproduce the kind of traffic real users actually generate. But if one of the eventual bottlenecks was literally a slow network link, then a much simpler high-volume saturation test might have found it earlier. Just push enough traffic through each part of the system and see where the throughput stops increasing.
Realistic load tells you what breaks when customers use the site normally, just at scale. A saturation test tells you where the hard capacity limits are. Those are different questions, and in hindsight we should have been asking both.