About the ClusterGroupUpgrade CR
The Topology Aware Lifecycle Manager (TALM) builds the remediation plan from the ClusterGroupUpgrade CR for a group of clusters. You can define the following specifications in a ClusterGroupUpgrade CR:
-
Clusters in the group
-
Blocking
ClusterGroupUpgradeCRs -
Applicable list of managed policies
-
Number of concurrent updates
-
Applicable canary updates
-
Actions to perform before and after the update
-
Update timing
You can control the start time of an update using the enable field in the ClusterGroupUpgrade CR.
For example, if you have a scheduled maintenance window of four hours, you can prepare a ClusterGroupUpgrade CR with the enable field set to false.
You can set the timeout by configuring the spec.remediationStrategy.timeout setting as follows:
spec
remediationStrategy:
maxConcurrency: 1
timeout: 240
You can use the batchTimeoutAction to determine what happens if an update fails for a cluster.
You can specify continue to skip the failing cluster and continue to upgrade other clusters, or abort to stop policy remediation for all clusters.
Once the timeout elapses, TALM removes all enforce policies to ensure that no further updates are made to clusters.
To apply the changes, you set the enabled field to true.
For more information see the "Applying update policies to managed clusters" section.
As TALM works through remediation of the policies to the specified clusters, the ClusterGroupUpgrade CR can report true or false statuses for a number of conditions.
|
|
After TALM completes a cluster update, the cluster does not update again under the control of the same
|