samba-mirror

mirror of https://github.com/samba-team/samba.git synced 2025-01-12 09:18:10 +03:00

Author	SHA1	Message	Date
Martin Schwenke	174449c1e0	ctdb-recoverd: Release recovery lock on exit The recovery lock helper must exit when it notices its parent is gone. However, that can take a few seconds. The usual way of terminating the recovery daemon is for the main ctdbd to send it a SIGTERM. Installing a handler is nice and simple. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-06-08 00:51:29 +02:00
Martin Schwenke	75717ac667	ctdb-recoverd: Add handler for lost recovery lock If the process holding the recovery lock terminates unexpectedly then the recovery daemon needs to know that the lock is no longer held. While here, rename hold_reclock_handler() to take_reclock_handler() so there is a clear difference between the two handler names. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-06-08 00:51:29 +02:00
Martin Schwenke	95a7920d22	ctdb-cluster-mutex: Register an extra handler for when mutex is lost Pass NULL if not needed. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-06-08 00:51:29 +02:00
Martin Schwenke	4f0ca0107c	ctdb-cluster-mutex: ctdb_cluster_mutex() registers handler and private data Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-06-08 00:51:29 +02:00
Martin Schwenke	145ddcbe37	ctdb-cluster-mutex: Drop cluster_mutex_handler() ctdb and handle arguments This makes the API more general. If they are needed in a handler then they can be in the private data. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-06-08 00:51:29 +02:00
Martin Schwenke	a192364a12	ctdb-recoverd: Simplify reclock handler Do the interesting work outside the handler. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-06-08 00:51:29 +02:00
Martin Schwenke	197264dfe7	ctdb-recoverd: Recovery lock handle should be in recovery deamon context This shouldn't be in the CTDB context. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-06-08 00:51:29 +02:00
Martin Schwenke	5c4744e69d	ctdb-cluster-mutex: Pass a talloc context to allocate the handle off Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-06-08 00:51:28 +02:00
Martin Schwenke	58be187de0	ctdb-recoverd: No need to reset reclock handler It won't be called more than once by the cluster mutex code. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-06-08 00:51:28 +02:00
Martin Schwenke	630f169653	ctdb-recoverd: Fix buggy function return on memory allocation failure Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-06-08 00:51:28 +02:00
Martin Schwenke	dbd4e67aee	ctdb-recoverd: Don't expose internal cluster mutex status Just expose whether the lock was taken. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-06-08 00:51:28 +02:00
Martin Schwenke	fdd214ce6a	ctdb-daemon: Rename recovery lock file to just recovery lock It isn't necessarily a file. Don't bother changing the control, since it doesn't pervade the code. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-06-08 00:51:28 +02:00
Martin Schwenke	1127f3ae1e	ctdb-recovery: Don't update recovery lock from daemon It can't change after startup. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-06-08 00:51:28 +02:00
Martin Schwenke	23823f128f	ctdb-recovery: Don't sync recovery lock across cluster Support for updating the recovery lock is being removed because it isn't possible to recover from failure. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-06-08 00:51:28 +02:00
Amitay Isaacs	f8141e91a6	ctdb-recoverd: Freeze databases whenever the node is INACTIVE If the node becomes stopped or banned after recovery is marked active, then it will never freeze the databases, and hence the node will keep banning itself indefinitely, until ctdbd is restarted. This is a regression from 4.3, introduced with `b4357a79d9` and `d8f3b490bb` BUG: https://bugzilla.samba.org/show_bug.cgi?id=11945 Signed-off-by: Amitay Isaacs <amitay@gmail.com> Reviewed-by: Michael Adam <obnox@samba.org> Reviewed-by: Martin Schwenke <martin@meltin.net> Autobuild-User(master): Michael Adam <obnox@samba.org> Autobuild-Date(master): Wed Jun 1 17:36:12 CEST 2016 on sn-devel-144	2016-06-01 17:36:12 +02:00
Martin Schwenke	f9d4cb4c29	ctdb-recoverd: Unify takeover run triggering code in main loop Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com> Autobuild-User(master): Amitay Isaacs <amitay@samba.org> Autobuild-Date(master): Fri May 13 17:15:57 CEST 2016 on sn-devel-144	2016-05-13 17:15:57 +02:00
Martin Schwenke	e3e4f37c41	ctdb-recoverd: Add early return in srvid_requests_reply() Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-05-13 13:47:17 +02:00
Martin Schwenke	ebbeab74ed	ctdb-recoverd: Drop an unnecessary log message do_takeover_run() will logs something at NOTICE level anyway. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-05-13 13:47:17 +02:00
Martin Schwenke	2a93b8423b	ctdb-recoverd: Move takeover run checks after recover checks If a recovery is going to be done then this will be followed by a takeover run anyway. So, there's no use doing the takeover run checks, potentially doing a takeover run and then doing a recovery. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-05-13 13:47:17 +02:00
Martin Schwenke	662f06de9f	ctdb-recoverd: Drop explicit check to flag takeover run needed The recovery daemon should be less involved in the service monitoring logic. The cases handled here are already handled elsewhere: * When a node becomes unhealthy/healthy the monitoring code will trigger a takeover run * When a node is disabled/enabled the ctdb CLI tool will trigger a takeover run Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-05-13 13:47:17 +02:00
Martin Schwenke	9dc3b117e2	ctdb-takeover: Recovery daemon no longer passes fail callback Banning is now handled by the takeover code sending banning credit messages. This commit makes a change in behaviour quite obvious. Takeover runs were initiated from several locations in the code but banning was only done from one of these locations. Now banning can be done from any failed takeover run. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-05-13 13:47:17 +02:00
Martin Schwenke	866ca591d4	ctdb-recoverd: Fold IP allocation house-keeping into IP verification Now all the IP takeover code for non-master node is in this function. The function can always be renamed to something more suitable. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com> Autobuild-User(master): Martin Schwenke <martins@samba.org> Autobuild-Date(master): Fri May 6 15:10:59 CEST 2016 on sn-devel-144	2016-05-06 15:10:59 +02:00
Martin Schwenke	4947789b2a	ctdb-recoverd: Clean up local IP verification Update log levels and messages, comments and wrapping of long lines. No functional changes. Note that interfaces_have_changed() already does adequate logging. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-05-06 11:39:09 +02:00
Martin Schwenke	bdcc796f3c	ctdb-recoverd: Skip known IP address checking when it is disabled When public IP checking is disabled, verify_local_ip_allocation() still retrieves known IP addresses and runs through a loop that does nothing. Instead, completely skip the retrieval and checking loop. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-05-06 11:39:09 +02:00
Martin Schwenke	fc4cbf5528	ctdb-recoverd: Check that IP failover is active in IP verification This makes verify_local_ip_allocation() self-contained and simplifies main_loop(). Due to indentation changes, this commit is most easily read when ignoring whitespace. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-05-06 11:39:09 +02:00
Martin Schwenke	ff28cbb73d	ctdb-recoverd: Call election when necessary in recovery master validation There is no need to return one of several states and then trigger an election for one of those return states. Have the recovery master validation trigger the election directly and just return whether monitoring should continue. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-05-06 11:39:09 +02:00
Martin Schwenke	e8c33aa24a	ctdb-recoverd: Simplify return values when updating local flags Change this to return just 0 or -1. It isn't monitoring anything. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-05-06 11:39:09 +02:00
Martin Schwenke	0a9401ff0e	ctdb-recoverd: Drop unreachable code update_local_flags() never returns MONITOR_ELECTION_NEEDED, so drop this entire if-statement. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-05-06 11:39:09 +02:00
Martin Schwenke	721f64511c	ctdb-recovery: Move recovery lock latency updating to handler The cluster mutex code already passes the latency and expects the handler to update the statistics. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-04-28 09:39:17 +02:00
Martin Schwenke	bcb838ba1e	ctdb-recovery: Move recovery lock functions to recovery daemon code ctdb_recovery_have_lock(), ctdb_recovery_lock(), ctdb_recovery_unlock() are only used by recovery daemon, so move them there. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-04-28 09:39:17 +02:00
Amitay Isaacs	ae366fb932	ctdb-recoverd: Add message handler to assigning banning credits This will be called from recovery helper to assign banning credits to misbehaving node. Signed-off-by: Amitay Isaacs <amitay@gmail.com> Reviewed-by: Martin Schwenke <martin@meltin.net>	2016-03-25 03:26:16 +01:00
Martin Schwenke	c9e69a4b2e	ctdb-recoverd: Drop use of DeferredRebalanceOnNodeAdd tunable If set, this was used to setup an IP takeover run on a timer after certain updates to the public IP address configuration (e.g. "ctdb addip"). However, "ctdb reloadips" completely manages public IP reconfiguration and avoids the anomalies that DeferredRebalanceOnNodeAdd was introduced to work around. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2016-03-10 03:34:19 +01:00
Amitay Isaacs	19a411f839	ctdb-recovery: Create recovery databases in state dir This matches the behaviour during serial database recovery. Signed-off-by: Amitay Isaacs <amitay@gmail.com> Reviewed-by: Martin Schwenke <martin@meltin.net> Autobuild-User(master): Martin Schwenke <martins@samba.org> Autobuild-Date(master): Thu Feb 11 08:01:14 CET 2016 on sn-devel-144	2016-02-11 08:01:14 +01:00
Martin Schwenke	56ce230de7	ctdb-recoverd: Fix some uninitialised memory issues The first element of these structures is a 32-bit PNN. On 64-bit systems this field can be followed by 32-bits of padding. When the structures are copied this can cause uninitialised memory to be copied. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Michael Adam <obnox@samba.org>	2016-01-12 19:16:17 +01:00
Martin Schwenke	bd7c94d5ac	ctdb-recoverd: Drop function unban_all_nodes() It hasn't worked since commit `cda5f02c7c` in 2009, which reworked the banning code. Since then ctdb_control_modflags() has contained a comment saying: /* we don't let other nodes modify our BANNED status */ Unbanning all nodes originally occurred here when the recovery master role moved to a new node. The logic could have been meant for the case when the old recovery master was malfunctioning, so got banned. If any other nodes had been banned by this recovery master then they would be unbanned. However, this would also unban the old recovery master, which is probably suboptimal. The logic would also trigger if a node was banned for a good reason and then the recovery master was stopped. So, apart from doing nothing, the logic is too simplistic so might as well be removed. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2015-12-04 09:17:17 +01:00
Christof Schmitt	03b27bd139	ctdb: Use prctl_set_comment from lib/util Signed-off-by: Christof Schmitt <cs@samba.org> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2015-11-18 04:05:13 +01:00
Martin Schwenke	44bf7c2a12	ctdb-recoverd: Factor out recovery master validation Starting to untangle cluster management, database recovery and public IP allocation. This is a non-trivial subset of the cluster management code that runs in the recovery daemon on all nodes. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com> Autobuild-User(master): Amitay Isaacs <amitay@samba.org> Autobuild-Date(master): Mon Nov 16 11:47:45 CET 2015 on sn-devel-104	2015-11-16 11:47:44 +01:00
Martin Schwenke	e44957fc8b	ctdb-recmaster: Update capabilities before calling first election Capabilities are used when computing an election result so having them up-to-date seems like a good idea. Also update several instances of an ambiguous comment. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2015-11-16 08:42:12 +01:00
Martin Schwenke	c5e50a474b	ctdb-recoverd: Move VNN map retrieval to where it is needed The VNN map is only needed on the recovery master, so no need for all recovery daemons to retrieve it. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2015-11-16 08:42:12 +01:00
Martin Schwenke	d1f996a50f	ctdb-recoverd: Drop explicit check for recovery lock This is already handled in update_recovery_lock(), which is called immediately before. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2015-11-16 08:42:12 +01:00
Martin Schwenke	1499f3e301	ctdb-recoverd: Simplify using TALLOC_FREE() The only non-obvious part here is dropping the setting of the nodemap local variable to NULL. If the following control succeeds then it is set, otherwise return and it doesn't matter. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2015-11-16 08:42:12 +01:00
Martin Schwenke	050e64b647	ctdb-recoverd: Clarify that recmaster is being set on the current node That is, using CTDB_CURRENT_NODE makes this more obvious. Also fix incorrect error messages. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2015-11-16 08:42:12 +01:00
Martin Schwenke	0833e478c3	ctdb-recoverd: Do not sanity check recovery master with local daemon Each recovery daemon knows who the recmaster is and is in sync with its local daemon. The recovery master is running this check so do not bother checking with its local daemon - both agree that it is the recovery master. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2015-11-16 08:42:12 +01:00
Martin Schwenke	d8decd0b1d	ctdb-recoverd: Don't retrieve recovery master from local daemon The recovery daemon already knows which node is the master. This relies on rec->recmaster being correctly initialised and correctly set during elections. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2015-11-16 08:42:12 +01:00
Martin Schwenke	e90cab7073	ctdb-recoverd: Explicitly set initial recovery master to unknown Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2015-11-16 08:42:11 +01:00
Martin Schwenke	018077f3b0	ctdb-recoverd: Do not set recovery master during recovery Recovery should not do cluster management functions. Setting the recovery master should only be done via an election. Main loop will determine if recovery master is inconsistent across the cluster and force an election if necessary. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2015-11-16 08:42:11 +01:00
Martin Schwenke	4b37cc7cf6	ctdb-recoverd: Have recovery daemon remember election result The recovery daemon pushes knowledge of recovery master election progress/result to local daemon. It then retrieves that information again. Instead, have the recovery daemon reliably track election progress/result in rec->recmaster so it doesn't need to be retrieved. Be careful to maintain consistency by only doing this when the local daemon has been updated. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2015-11-16 08:42:11 +01:00
Martin Schwenke	6f8837528f	ctdb-recoverd: Clarify recovery master validation logic There can be no holes in the nodemap. Even if a node has been deleted it will take a slot in the nodemap. The only exception is that the nodemap shrinks if nodes are deleted from the end. That should never include the master because a node should be shutdown before being deleted, and an election should already have take place. To avoid walking off the end of the nodemap nodes array just confirm that the master node's PNN is a valid index into the array. No need to walk through the nodemap. After this, in this section of the code j is now invalid. So use the master's PNN to index into the nodemap. This is safe. In the process, clean up some log messages to avoid saying "Force reelection". It's just an "election". Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com>	2015-11-16 08:42:11 +01:00
Amitay Isaacs	f50db5cba5	ctdb-server: Replace ctdb_logging.h with common/logging.h Signed-off-by: Amitay Isaacs <amitay@gmail.com> Reviewed-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Michael Adam <obnox@samba.org>	2015-11-16 00:46:15 +01:00
Martin Schwenke	0886637a5e	ctdb-recoverd: Reload remote IPs as part of takeover run This is currently done before each IP takeover run, so just factor it in. ctdb_reload_remote_public_ips() becomes static. Signed-off-by: Martin Schwenke <martin@meltin.net> Reviewed-by: Amitay Isaacs <amitay@gmail.com> Autobuild-User(master): Amitay Isaacs <amitay@samba.org> Autobuild-Date(master): Thu Nov 12 09:28:45 CET 2015 on sn-devel-104	2015-11-12 09:28:45 +01:00

1 2 3 4 5 ...

456 Commits