HA setup for clusters with regular maintenance windows
Overview
In some clusters, nodes are rotated regularly, for example during node image upgrades or maintenance windows in which nodes are drained and restarted one after another. In such clusters the HDFS storage of the SUSE® Observability topology database is restarted many times within a short period.
The 150-ha, 250-ha and 500-ha sizing profiles run HDFS with 3 data nodes and store every block twice. When data nodes restart in quick succession while data is being written, HDFS becomes more vulnerable to further node failures. While a data node restarts, at most 2 data nodes are active, so any additional node issue can cause downtime at best, or data corruption or data loss at worst. For clusters with regular maintenance windows, we recommend using the HDFS settings of the 4000-ha profile, which store every block 3 times on 5 data nodes:
150-ha, 250-ha, 500-ha |
Recommended (as in 4000-ha) |
|
|---|---|---|
HDFS data nodes |
3 |
5 |
Copies of every block ( |
2 |
3 |
Minimum copies of a block ( |
2 |
2 |
These settings add resilience, but they do not replace a careful maintenance procedure. For general guidance on maintaining and rotating nodes, including the checks to perform before disrupting the next node, see Maintain or rotate nodes. That section is part of the Longhorn page, but it applies regardless of Kubernetes distribution or cluster management platform.
Requirements
-
SUSE® Observability Helm chart version 2.10.0 or later, installed with the
150-ha,250-haor500-hasizing profile. -
At least 5 Kubernetes nodes. Every HDFS data node runs on a separate node.
-
Capacity for 2 additional HDFS data nodes. Each data node requests 600m CPU and 4Gi memory and uses its own persistent volume (250Gi by default). Because every block is stored 3 times across 5 data nodes, the amount of data per data node does not increase.
Configure the HDFS settings
Add the following to the values.yaml file that you use to install or upgrade SUSE® Observability:
hbase:
hdfs:
replication: 3
minReplication: 2
datanode:
replicaCount: 5
Apply the configuration by installing or upgrading SUSE® Observability:
helm upgrade \
--install \
--namespace suse-observability \
--values values.yaml \
suse-observability \
suse-observability/suse-observability
If you deploy with --set flags instead of a values file, add these flags to your helm upgrade command:
--set hbase.hdfs.replication=3 \
--set hbase.hdfs.minReplication=2 \
--set hbase.hdfs.datanode.replicaCount=5 \
On an existing installation, the new settings apply to data written after the upgrade. Data that already exists keeps 2 copies until it is rewritten.