Hello, Snap! Congratulations on a successful IPO! You made the right choice by choosing multiple cloud service providers (GCP[1] and AWS[2]) for your infrastructure needs. Here is my recommendation on why you should also consider an on-prem alternative such as OpenStack or Cloud Foundry. Please note that this has got nothing to do with the recent outages, we all know that failure is inevitable. Please also note that operating on-prem infrastructure is hard. This recommendation is only based on my understanding of Snapchat’s workloads and a top-of-the-envelope TCO analysis based on such workloads. You may also want to consider TCO/ ROI analyses from other sources such as existing customers (Snapdeal using OpenStack), or analysts (IDC on Red Hat OpenShift).

Estimated Compute & Storage Requirements

Since the content (images/ videos/ stories) that your users share is ephemeral (that they are deleted within 24 hours), you have the unique advantage of having low data gravity. However, you are not completely out of it. For example, you need to manage your user account details. You also need to store and manage data that will be used towards your advertising revenue. You are also growing at an incredible rate across multiple geographic locations.

Based on your existing usage patterns as cited in your S1 filing[1] and your projected growth, here are interesting numbers on the volume of data shared and stored by Snapchat users:

Active Users Per Day: 150+ M

‘Snaps’ per day: 2.5 B

Stories Shared: 25% of daily users post stories of an average duration of 20–30 minutes each.

Assuming each ‘Snap’ is about 15 KB [3], that is more than 30 TB of data shared as snaps every day. Assuming each story is about 1 MB per minute, about 750 TB — 1.1 PB of data every day as stories (I am skeptical about the number and size of stories shared, though. I would expect them to be an order smaller currently).

If you continue your current rate of growth before the inevitable plateau, your users will be potentially be sharing multiple petabytes of data in a day. Since you delete snaps and stories within 24 hours, your maximum size required for storage may not exceed 10 PB on any given day, though. It is also to be noted that since your users are across multiple geographic locations and not all stories/ snaps need to be stored in same geo location. Given the short lifetimes of your content and heavy transfer rates, your data transfer and bandwidth costs might be higher than the storage costs themselves.

Finding your compute requirements is bit tricky. Extrapolating backward from the GCP TCO Calculator [4] and assuming that part of (10%) of the total amount of spend ($400 M per year) you’ve committed to GCP is to be spent on compute, you seem to be using over 200K cores! While it is entirely possible you utilize these many cores, we can safely assume that I am off by an order.

Based on above assumptions, I expect one of your regions/ locations to have following compute and storage requirements:

# of CPU Cores: 20,000 (for compute)

Object Storage: 1 PB (for ephemeral storage)

Block Storage: 10 TB (for user data, accounts, OS, etc)

AWS vs On-Prem TCO Comparisons

I used AWS TCO Calculator [5] to compare if on-prem options would provide you any cost savings. I would have preferred to use GCP TCO calculator, GCP being your primary cloud service provider. But it doesn’t support additional options such as storage and networking as of now. It also only compares with AWS with no comparison with on-prem options. I used following additional numbers for on-prem options

Average IT Admin Salary: $140,000 per year (based on US average for OpenStack admins; VMware and Microsoft admin averages are much lower)

Number of VMs managed per admin: 400 (based on the standard that AWS uses for its own TCO)

# of CPU Cores = 2 per VM

Memory = 4 GB per VM

Datacenter Network Bandwidth:1 Gbps

Average Utilization :90% per VM

Object Store Access Frequency: 100% as they are deleted with in 24 hours

% Change in Block Storage: 50% (needed to provide business continuity — could be higher for you)

Hypervisor: KVM/ Xen

Let us also use three different geographic regions considering your user base across multiple geographic locations. Here are the comparisons based on regions — AWS US West (Oregon), AWS EU (Ireland) and AWS Asia Pacific (Singapore).

Figure 1 AWS TCO Comparison — US West (Oregon)

Figure 2 AWS TCO Comparison — EU (Ireland)

Figure 3 AWS TCO Comparison — AWS Asia Pacific (Singapore)

You can notice that compute is the largest component of the spend, with IT-Labor being not an insignificant component. Major cost savings seem to be from storage related costs, as I had indicated earlier. You may also notice that if you use bulkier virtual machines with larger memory and more cores, AWS could provide better TCO. AWS TCO also includes business level support which on-prem TCO calculations shown may not include completely. These calculations also don’t include any special discounts that AWS might offer.

Though these comparisons are just indicative, it appears that you will gain cost benefits by considering on-prem alternatives, given your scale.

On-prem options such as OpenStack also provides you the ability to treat your infrastructure as code, thereby making operations across bare-metal, virtual machines and containers easier. With OpenStack, you are also positioning yourself better for potential re-partitioning of your workloads in the future.

You may also want to consider the efforts needed to operate at scale using cloud platforms. To be more successful and efficient using GCP, you probably have developed enough automation in your deployment and operations. You need to develop same amount of sophistication on the next platform you choose (AWS). Alternatives such as Cloud Foundry help you to minimize such duplicate automation efforts by supporting multi-cloud options.

Finally, by leveraging on-prem options, you are also minimizing the risks you have cited in your S1 filing related to the dependency on other cloud service providers.

Summary

From your S1[2]filing statement, it appears that you are not opposed to on-prem alternatives. This is my humble attempt to quantify the benefits of on-prem options for users of your scale. It is not going to be easy, but there is plenty of help available should you need.

Congrats again on a successful IPO and wishing you the best!

References

  1. Snap Inc. S1 Filing: https://www.sec.gov/Archives/edgar/data/1564408/000119312517029199/d270216ds1.htm

  2. Snap Inc. S1 Filing Addendum: https://www.sec.gov/Archives/edgar/data/1564408/000119312517034993/d270216ds1a.htm

  3. Snapchat file sizes: http://www.ibtimes.com/heres-crazy-amount-cellular-data-snapchat-consumes-how-stop-it-1938313

  4. GCP TCO Calculator: https://cloud.google.com/pricing/tco/

  5. AWS TCO Calculator: https://awstcocalculator.com/#