Replication or cluster – which is better?

If you have come here, you are most likely looking for an answer to the question: which is better – replication or clustering?…

If you have come here, you are most likely looking for an answer to the question: which is better – replication or clustering? It is difficult to provide a clear-cut answer without knowing the specifics of the business and the client’s needs. However, one thing can be said for certain. Both solutions, when combined, provide the highest level of availability, although this is also the most expensive option, which few companies choose. Most micro and SME companies do not have adequately secured infrastructure and only start looking for ways to prevent further failures after the first (unfortunately) serious outage. Rarely, but it also happens that a company turns out to be “resistant” to a prolonged downtime of IT systems and does not need a highly available environment. In this article, we will therefore answer the question: which is better, replication or clustering? We encourage you to read on.

What is replication?

Replication is a copy of a database or set of data that exists alongside the original data source. Replication can be used for various purposes, including increasing availability, resilience to failures, and system performance. The replication process involves continuously copying data from one location (the primary) to one or more other locations (replicas). Replicas can be used to distribute query workloads across different servers, enabling faster query processing and reducing response times. Depending on the chosen strategy, replication can be synchronous, where data is replicated immediately, or asynchronous, which allows for some delay in data propagation.

What is a cluster?

A cluster is a group of interconnected servers working together as a single system. Clustering aims to increase availability and performance by distributing workloads and providing redundancy. In a cluster, tasks and resources are managed collectively, enabling more efficient resource utilization and better scalability. A cluster can be used in various applications, from scientific computing to processing large datasets. A key aspect of clusters is their ability to automatically detect and manage failures, which helps minimize downtime and ensure operational continuity. Clusters can be configured in various ways, depending on the specific needs and requirements of the application.

For critical systems, a cluster is always worth having, but is it really?

Using a two-node cluster + data storage is probably the most popular method of building a highly available environment. However, a limited budget often simply does not allow for the purchase of two data storage systems, so a company or institution decides to purchase a single storage array.

  • Zero Downtime – in the event of a failure of one host, the other takes over its role, and if several TCP packets are lost, the failover will occur automatically.
  • Flexibility – host updates can be performed while the other host is 100% active.

Only two advantages? In my opinion, that’s where the list ends. So what are the disadvantages of a two-node cluster with a single storage array? Let’s take a look.

  • “Single Point of Failure” – a single storage array is unfortunately very often the most common SPoF in this type of architecture.
  • Costs. If a storage array with 10K SAS HDDs is sufficient, you probably won’t need to invest a lot. However, an array with SAS SSDs or U.3 NVMe drives is a much greater expense. In addition, to fully utilize NVMe drives, you should consider an NVMe-oF network interface.
  • Testing automatic failover, fault tolerance, and correct network configuration is practically impossible with a single data storage system, as we cannot disconnect it from production even temporarily.

In my opinion, the arguments above represent the main disadvantages versus advantages of a two-node cluster solution.

Many companies have already experienced the impact that a “Single Point of Failure” can have on business continuity. Unfortunately, it was often only after unpleasant incidents that they implemented a Disaster Recovery Plan aimed at improving the situation and properly securing their infrastructure. Remember not to repeat the mistakes of others and be prepared!

What happened? HP’s SSD drive failure

HP SAS drives suddenly stopped working. It turned out that the drives contained faulty firmware, which caused them to fail after reaching a specific operating-hours limit.

There was nothing that could be done. The only solution was to reconfigure the storage array and restore the entire environment from a backup. It took some time, but the recovery was successful and the client recovered the entire environment. However, they lost a full day of work because backups of all systems were performed only once a day.

There have probably been many similar incidents. So how can you safely and efficiently update drive firmware if you have only one storage system? And what happens if it doesn’t start afterward? How will we recover our infrastructure at 10 p.m. on a Saturday evening?

I have many examples like this, but this was the most serious one. In theory, the user was in no way at fault. I say “in theory” because HP had been sending error notifications, but let’s be honest – who actually read them?

So, is replication better than a cluster?

Comparing it to a cluster, let’s take a look at the advantages of replication:

  1. The number one advantage for me is having a separate data storage system. This is a key factor and one of the main reasons why it is worth considering implementing replication.
  2. Testing is much easier. During testing, we have an environment running the same version of the system.
  3. Simpler administration.
  4. Easy replication monitoring.
  5. Lower purchase and maintenance costs. You don’t need SAS drives or Microsoft licenses for a second host. The only exception is if you don’t already have the necessary licenses under the SPLA model or don’t have Software Assurance.
  6. The ability to perform synchronous replication, meaning replication without data loss.
  7. The ability to perform asynchronous replication. This allows us to replicate data to the other side of the world without any problems over a WAN connection.
bezpieczenstwo it

Those are probably all the main advantages. Let’s move on to the disadvantages.

  1. Synchronous replication is expensive. Not all hypervisors support it, and it may require additional licenses.
  2. Asynchronous replication can result in data loss – this is its main disadvantage.
  3. Application servers do not switch over automatically.
  4. Not all applications support replication.

When is asynchronous replication the best choice?

  • The servers have been encrypted. We should have access to a second replica, provided that the network has been properly configured and checkpoints have been added to the configuration.
  • We need to test something on a server with a production configuration. We can temporarily shut it down, start the replica on a private network, and perform the necessary tests.

REMEMBER! Do not treat replication as a backup!

it centrum

How to choose between replication and clustering?

The decision between replication and clustering should be based on a detailed analysis of system requirements, availability, and fault tolerance, all of which are crucial to an organization’s operations. Before making a decision, it is important to understand the purpose of each solution and its impact on the IT infrastructure.

When choosing between replication and clustering, the following aspects should be taken into account:

  1. Availability requirements – it is essential to determine how important continuous availability of data and services is to the company’s operations.
  2. Performance requirements – analyzing workloads and throughput can indicate whether horizontal scalability, which is offered by clustering, is more important.
  3. Budget – the costs of implementing and maintaining a cluster are generally higher than those of simple replication solutions, which may be a deciding factor depending on the available financial resources.
  4. Administrative complexity – managing clusters requires more advanced skills and tools than managing individual replicas, which can influence decisions regarding human resources.

In conclusion

Based on my experience, asynchronous replication will be sufficient for most customers. For 90% of companies, the most critical system is cloud-based email, while the remaining 10% are manufacturing companies that need to accurately estimate the potential cost of data loss and downtime.

Defining the RTO and RPO metrics is the foundation for choosing the right solution. The ideal approach is to combine both technologies, but not every organization can afford such an investment.

If you’re still wondering which solution to choose, I encourage you to take a look at our offering and get in touch. Together, we can select the right solution for your company.

Read more

Related articles

View all articles