| Red Hat Cluster Manager: The Red Hat Cluster Manager Installation and Administration Guide | ||
|---|---|---|
| Prev | Appendix B. Supplementary Software Information | Next |
Understanding cluster behavior when significant events occur can assist in the proper management of a cluster. Note that cluster behavior depends on whether power switches are employed in the configuration. Power switches enable the cluster to maintain complete data integrity under all failure conditions.
The following sections describe how the system will respond to various failure and error scenarios.
In a cluster configuration that uses power switches, if a system hangs, the cluster behaves as follows:
The functional cluster system detects that the hung cluster system is not updating its timestamp on the quorum partitions and is not communicating over the heartbeat channels.
The functional cluster system power-cycles the hung system. Alternatively, if watchdog timers are in use, a failed system will reboot itself.
The functional cluster system restarts any services that were running on the hung system.
If the previously hung system reboots, and can join the cluster (that is, the system can write to both quorum partitions), services are re-balanced across the member systems, according to each service's placement policy.
In a cluster configuration that does not use power switches, if a system hangs, the cluster behaves as follows:
The functional cluster system detects that the hung cluster system is not updating its timestamp on the quorum partitions and is not communicating over the heartbeat channels.
Optionally, if watchdog timers are used, the failed system will reboot itself.
The functional cluster system sets the status of the hung system to DOWN on the quorum partitions, and then restarts the hung system's services.
If the hung system becomes active, it notices that its status is DOWN, and initiates a system reboot.
If the system remains hung, manually power-cycle the hung system in order for it to resume cluster operation.
If the previously hung system reboots, and can join the cluster, services are re-balanced across the member systems, according to each service's placement policy.
A system panic (crash) is a controlled response to a software-detected error. A panic attempts to return the system to a consistent state by shutting down the system. If a cluster system panics, the following occurs:
The functional cluster system detects that the cluster system that is experiencing the panic is not updating its timestamp on the quorum partitions and is not communicating over the heartbeat channels.
The cluster system that is experiencing the panic initiates a system shut down and reboot.
If power switches are used, the functional cluster system power-cycles the cluster system that is experiencing the panic.
The functional cluster system restarts any services that were running on the system that experienced the panic.
When the system that experienced the panic reboots, and can join the cluster (that is, the system can write to both quorum partitions), services are re-balanced across the member systems, according to each service's placement policy.
Inaccessible quorum partitions can be caused by the failure of a SCSI (or Fibre Channel) adapter that is connected to the shared disk storage, or by a SCSI cable becoming disconnected to the shared disk storage. If one of these conditions occurs, and the SCSI bus remains terminated, the cluster behaves as follows:
The cluster system with the inaccessible quorum partitions notices that it cannot update its timestamp on the quorum partitions and initiates a reboot.
If the cluster configuration includes power switches, the functional cluster system power-cycles the rebooting system.
The functional cluster system restarts any services that were running on the system with the inaccessible quorum partitions.
If the cluster system reboots, and can join the cluster (that is, the system can write to both quorum partitions), services are re-balanced across the member systems, according to each service's placement policy.
A total network connection failure occurs when all the heartbeat network connections between the systems fail. This can be caused by one of the following:
All the heartbeat network cables are disconnected from a system.
All the serial connections and network interfaces used for heartbeat communication fail.
If a total network connection failure occurs, both systems detect the problem, but they also detect that the SCSI disk connections are still active. Therefore, services remain running on the systems and are not interrupted.
If a total network connection failure occurs, diagnose the problem and then do one of the following:
If the problem affects only one cluster system, relocate its services to the other system. Then, correct the problem and relocate the services back to the original system.
Manually stop the services on one cluster system. In this case, services do not automatically fail over to the other system. Instead, restart the services manually on the other system. After the problem is corrected, it is possible to re-balance the services across the systems.
Shut down one cluster system. In this case, the following occurs:
Services are stopped on the cluster system that is shut down.
The remaining cluster system detects that the system is being shut down.
Any services that were running on the system that was shut down are restarted on the remaining cluster system.
If the system reboots, and can join the cluster (that is, the system can write to both quorum partitions), services are re-balanced across the member systems, according to each service's placement policy.
If a query to a remote power switch connection fails, but both systems continue to have power, there is no change in cluster behavior unless a cluster system attempts to use the failed remote power switch connection to power-cycle the other system. The power daemon will continually log high-priority messages indicating a power switch failure or a loss of connectivity to the power switch (for example, if a cable has been disconnected).
If a cluster system attempts to use a failed remote power switch, services running on the system that experienced the failure are stopped. However, to ensure data integrity, they are not failed over to the other cluster system. Instead, they remain stopped until the hardware failure is corrected.
If a quorum daemon fails on a cluster system, the system is no longer able to monitor the quorum partitions. If power switches are not used in the cluster, this error condition may result in services being run on more than one cluster system, which can cause data corruption.
If a quorum daemon fails, and power switches are used in the cluster, the following occurs:
The functional cluster system detects that the cluster system whose quorum daemon has failed is not updating its timestamp on the quorum partitions, although the system is still communicating over the heartbeat channels.
After a period of time, the functional cluster system power-cycles the cluster system whose quorum daemon has failed. Alternatively, if watchdog timers are in use, the failed system will reboot itself.
The functional cluster system restarts any services that were running on the cluster system whose quorum daemon has failed.
If the cluster system reboots and can join the cluster (that is, it can write to the quorum partitions), services are re-balanced across the member systems, according to each service's placement policy.
If a quorum daemon fails, and neither power switches nor watchdog timers are used in the cluster, the following occurs:
The functional cluster system detects that the cluster system whose quorum daemon has failed is not updating its timestamp on the quorum partitions, although the system is still communicating over the heartbeat channels.
The functional cluster system restarts any services that were running on the cluster system whose quorum daemon has failed. Under the unlikely event of catastrophic failure, both cluster systems may be running services simultaneously, which can cause data corruption.
If the heartbeat daemon fails on a cluster system, service failover time will increase because the quorum daemon cannot quickly determine the state of the other cluster system. By itself, a heartbeat daemon failure will not cause a service failover.
If the power daemon fails on a cluster system and the other cluster system experiences a severe failure (for example, a system panic), the cluster system will not be able to power-cycle the failed system. Instead, the cluster system will continue to run its services, and the services that were running on the failed system will not fail over. Cluster behavior is the same as for a remote power switch connection failure.
If the service manager daemon fails, services cannot be started or stopped until you restart the service manager daemon or reboot the system. The simplest way to restart the service manager is to first stop the cluster software and then restart it. For example, to stop the service, perform the following command:
/sbin/service cluster stop |
Then, to restart the cluster software, perform the following:
/sbin/service cluster start |