The Amazon S3 Glacier storage classes are purpose-built for data archiving, providing a low cost archive storage in the cloud. According to AWS, S3 Glacier storage classes provide virtually unlimited scalability and are designed for 99.999999999% (11 nines) of data durability, and they provide fast access to archive data and low cost.
$0
Per GB Per Month
Azure Data Lake Storage
Score 9.0 out of 10
N/A
Azure Data Lake Storage Gen2 is a highly scalable and cost-effective data lake solution for big data analytics. It combines the power of a high-performance file system with massive scale and economy to help you speed your time to insight. Data Lake Storage Gen2 extends Azure Blob Storage capabilities and is optimized for analytics workloads.
If your organization has a lot of archival data that it needs to be backed up for safekeeping, where it won't be touched except in a dire emergency, Amazon Glacier is perfect. In our case, we had a client that generates many TB of video and photo data at annual events and wanted to retain ALL of it, pre- and post- edit for potential use in a future museum. Using the Snowball device, we were able to move hundreds of TB of existing media data that was previously housed on multiple Thunderbolt drives, external RAIDs, etc, in an organized manner, to Amazon Glacier. Then, we were able to setup CloudBerry Backup on their production computers to continually backup any new media that they generated during their annual events.
Azure Data Lake is an absolutely essential piece of a modern data and analytics platform. Over the past 2 years, our usage of Azure Data Lake as a reporting source has continued to grow and far exceeds more traditional sources like MS SQL, Oracle, etc.
Since the rest of our infrastructure is in Amazon AWS, coding for sending data to Glacier just makes sense. The others are great as well, for their specific needs and uses, but having *another* third-party software to manage, be billed for, and learn/utilize can be costly in money and time.
Azure Data Lake Storage from a functionality perspective is a much easier solution to work with. It's implementation from Amazon EMR went smooth, and continued usage is definitely better. However, Amazon EMR was significantly cheaper overall between the high transaction fees and cost of storage due to growth. The two both have their advantages and disadvantages, but the functionality of Azure Data Lake Storage outweighed it's cost
We seldom need to access our data in Glacier; this means that it is a fraction of the cost of S3, including the infrequent-access storage class.
Transitioning data to Glacier is managed by AWS. We don't need our engineers to build or maintain log pipelines.
Configuring lifecycle policies for S3 and Glacier is simple; it takes our engineers very little time, and there is little risk of errant configuration.
Instead of having separate pools of storage for data we are now operating on a single layer platform which has cut down on time spent on maintaining those separate pools.
We have had more of an ROI with the scalability as we are able to control costs of storage when need be.
We are able to operate in a more streamlined approach as we are able to stay within the Azure suite of products and integrate seamlessly with the rest of the applications in our cloud-based infrastructure