The day when the whole world knows whether you have done your homework or not. When at worst you have to pull out the cringe pillow and survive, or proudly stand and enjoy Google Analytics.
Andrew Golrang is an industry figure who helped take Cervera to new omnichannel heights and sales records. He has just changed jobs and is now Product Director at Voyado but he remembers a Black Friday from long ago…
"Black Friday. The most fun day of the year. Over 6 months of planning. The warehouse is packed with goods. Thousands of items are about to find new homes. The email goes out. Tense wait. Orders start coming in. More and more. Every time you hit F5 a load of orders have dropped in. But then it stops. Strange. Where did the customers go? The website has gone down. Panic. Restart the web server. Hoping the customers come back. They do. It takes 15 minutes. The site dies again. Restart again. But then it sorts itself out. Success.
Black Friday has broken records every year. Every year we at Cervera pushed until we broke. And still the shopping day's success continued. It is amazing."
I think most of us who work with e-commerce can recognise ourselves. It really sends a shiver down your spine reading Andrew's description. For us, Black Friday has been many years of nervously watching response times in New Relic and switching over to the real-time figure in GA. Pestering the architect or the on-call developers with "is it holding?" while the client pings on Slack with "isn't it a bit slow?" or the even worse "are we down?". It is a stressful evening with a lot of nerves that, more often than not, ended up going well with sales records rather than tech panic.
Andrew continues:
"I really can miss the pulse Black Friday and other big shopping days gave in my new role. This Black Friday is the first time in 8 years I will hopefully have time to make a bargain or two myself, now that I am not personally responsible for any operational piece at either Cervera or Voyado. And from Voyado's perspective we are excited to the tips of our toes."
We at Commerce Mind recognise ourselves in Andrew's journey. We have all moved from working as vendors of e-commerce to clients and sitting with the responsibility for operations and everything working, to now working one step further away. But we have not stopped working on making sure the sites and apps still work.
The recipe is of course to prepare, and technically it usually starts already 6 months in advance with:
What has been done since the evaluation of the last Black Friday.
Mapping bottlenecks by looking at logs and running performance tests.
Picking low-hanging fruit (cutting silly database queries and making sure database indexes are right usually helps a lot).
Then come architectural changes, new systems, caching, entirely new solutions.
The hard part is that if you think about performance in your system last, it will not be easy to do much. There is no silver bullet. Performance is about the right architecture from the start, and after that it is many small streams.
If you have not started the performance work in good enough time, there are still things you can do to prepare in the weeks before, but the closer you get, the more compromises you have to be prepared to make, because less time means you may have to take shortcuts.
First and foremost, you need to figure out which parts of the system are more fragile than others. A good first step is to run a load test, which means simulating that thousands of visitors come to the site and do certain things at the same time. Building a good load test is an art in itself, because the test has to mimic real users' behaviour as much as possible. It is not uncommon to see a load test show fragility in one part of the system, but when Black Friday comes it turns out the load test was not representative and that fragility mattered less, while you completely missed another part that manages to bring the site down. That is why it is at least as important to look at your monitoring data from the peaks you have had recently, because that will show fragilities that real customers caused.
The two most important long-term tips Anders Ekdahl at Commerce Mind can share are:
Move logic out of your servers to as close to your customers as possible, something called Edge Compute. Here companies like Cloudflare and Fastly are at the strong forefront, but most CDN vendors have solutions for this. The more logic you move there, the more globally scalable you become. Feel free to contact us for a clearer walk-through of what this means and how you can build it up.
Make sure you have auto-scaling in as many parts of the system as possible. Even for the parts that say they have auto-scaling, you need to verify that it really works. It is not uncommon for auto-scaling to be there but to be too slow in practice and not cope with large and rapid increases. For the parts that do not have auto-scaling, you should design the systems landscape so that an increase in traffic does not linearly increase the pressure on those parts. This is where many of the advantages of headless and composable commerce lie, in building your system on top of such components. These are to a much greater extent built with auto-scaling and performance in mind.
Performance, scalability and reliability are things that must run through your whole systems landscape and architecture, and getting it in afterwards is anything but easy, which is a big reason why sites go down during Black Friday.
To be able to sleep well from Black Week through to January, you need to either buy an all-in-one solution where the platform guarantees scaling and performance. Or put together a best-of-breed architecture yourself where you pick out components with strong scalability. Which is right for you depends, but the best you can do is to talk to people who have made the same choice before.
Good luck today from all of us at Commerce Mind.