Adding and Removing Nodes From Subclusters
You will often want to add new nodes to and remove existing nodes from a subcluster. This ability lets you scale your database to respond to changing analytic needs.
Adding New Nodes to a Subcluster
You can add nodes to a subcluster to meet additional workloads. The nodes that you add to the subcluster must already be part of your cluster. These can be:
- Nodes that you removed from other subclusters.
- Nodes you added following the steps in Add Nodes to a Running AWS Cluster.
- Nodes you created using your cloud provider's interface, such as the AWS EC2 "Launch more like this" feature.
Do not add more nodes to a subcluster than you have shards in your database. Instead, create a new subcluster to hold the new nodes. Having more nodes in a subcluster than you have shards in the database is less efficient than having those nodes in a separate subcluser.
To add new nodes to a subcluster, use the db_add_node command of admintools:
$ adminTools -t db_add_node -h
Usage: db_add_node [options]
Options:
-h, --help show this help message and exit
-d DB, --database=DB Name of the database
-s HOSTS, --hosts=HOSTS
Comma separated list of hosts to add to database
-p DBPASSWORD, --password=DBPASSWORD
Database password in single quotes
-a AHOSTS, --add=AHOSTS
Comma separated list of hosts to add to database
-c SCNAME, --subcluster=SCNAME
Name of subcluster for the new node
--timeout=NONINTERACTIVE_TIMEOUT
set a timeout (in seconds) to wait for actions to
complete ('never') will wait forever (implicitly sets
-i)
-i, --noprompts do not stop and wait for user input(default false).
Setting this implies a timeout of 20 min.
--compat21 (deprecated) Use Vertica 2.1 method using node names
instead of hostnames
If you do not use the -c option, Vertica adds new nodes to the default subcluster (set to default_subcluster in new databases). This example adds a new node without specifying the subcluster:
$ adminTools -t db_add_node -p 'password' -d verticadb -s 10.11.12.117
Subcluster not specified, validating default subcluster
Nodes will be added to subcluster 'default_subcluster'
Verifying database connectivity...10.11.12.10
Eon database detected, creating new depot locations for newly added nodes
Creating depots for each node
Generating new configuration information and reloading spread
Replicating configuration to all nodes
Starting nodes
Starting nodes:
v_verticadb_node0004 (10.11.12.117)
Starting Vertica on all nodes. Please wait, databases with a
large catalog may take a while to initialize.
Checking database state
Node Status: v_verticadb_node0004: (DOWN)
Node Status: v_verticadb_node0004: (DOWN)
Node Status: v_verticadb_node0004: (DOWN)
Node Status: v_verticadb_node0004: (DOWN)
Node Status: v_verticadb_node0004: (UP)
Communal storage detected: syncing catalog
Multi-node DB add completed
Nodes added to verticadb successfully.
You will need to redesign your schema to take advantage of the new nodes.
To add nodes to a specific existing subcluster, use the db_add_node tool's -c option:
$ adminTools -t db_add_node -s 10.11.12.178 -d verticadb -p 'password' \
-c analytics_subcluster
Subcluster 'analytics_subcluster' specified, validating
Nodes will be added to subcluster 'analytics_subcluster'
Verifying database connectivity...10.11.12.10
Eon database detected, creating new depot locations for newly added nodes
Creating depots for each node
Generating new configuration information and reloading spread
Replicating configuration to all nodes
Starting nodes
Starting nodes:
v_verticadb_node0007 (10.11.12.178)
Starting Vertica on all nodes. Please wait, databases with a
large catalog may take a while to initialize.
Checking database state
Node Status: v_verticadb_node0007: (DOWN)
Node Status: v_verticadb_node0007: (DOWN)
Node Status: v_verticadb_node0007: (DOWN)
Node Status: v_verticadb_node0007: (DOWN)
Node Status: v_verticadb_node0007: (UP)
Communal storage detected: syncing catalog
Multi-node DB add completed
Nodes added to verticadb successfully.
You will need to redesign your schema to take advantage of the new nodes.
Updating Shard Subscriptions After Adding Nodes
After you add nodes to a subcluster they do not yet subscribe to shards. You can view the subscription status of all nodes in your database using the following query that joins the V_CATALOG.NODES and V_CATALOG.NODE_SUBSCRIPTIONS system tables:
=> SELECT subcluster_name, n.node_name, shard_name, subscription_state FROM v_catalog.nodes n LEFT JOIN v_catalog.node_subscriptions ns ON (n.node_name = ns.node_name) ORDER BY 1,2,3; subcluster_name | node_name | shard_name | subscription_state ----------------------+----------------------+-------------+-------------------- analytics_subcluster | v_verticadb_node0004 | | analytics_subcluster | v_verticadb_node0005 | | analytics_subcluster | v_verticadb_node0006 | | default_subcluster | v_verticadb_node0001 | replica | ACTIVE default_subcluster | v_verticadb_node0001 | segment0001 | ACTIVE default_subcluster | v_verticadb_node0001 | segment0003 | ACTIVE default_subcluster | v_verticadb_node0002 | replica | ACTIVE default_subcluster | v_verticadb_node0002 | segment0001 | ACTIVE default_subcluster | v_verticadb_node0002 | segment0002 | ACTIVE default_subcluster | v_verticadb_node0003 | replica | ACTIVE default_subcluster | v_verticadb_node0003 | segment0002 | ACTIVE default_subcluster | v_verticadb_node0003 | segment0003 | ACTIVE (12 rows)
You can see that none of the nodes in the newly-added analytics_subcluster have subscriptions.
To update the subscriptions for new nodes, call the REBALANCE_SHARDS function. You can limit the rebalance to the subcluster containing the new nodes by passing its name to the REBALANCE_SHARDS function call. The following example runs rebalance shards to update the analytics_subcluster's subscriptions:
=> SELECT REBALANCE_SHARDS('analytics_subcluster');
REBALANCE_SHARDS
-------------------
REBALANCED SHARDS
(1 row)
=> SELECT subcluster_name, n.node_name, shard_name, subscription_state FROM
v_catalog.nodes n LEFT JOIN v_catalog.node_subscriptions ns ON (n.node_name
= ns.node_name) ORDER BY 1,2,3;
subcluster_name | node_name | shard_name | subscription_state
----------------------+----------------------+-------------+--------------------
analytics_subcluster | v_verticadb_node0004 | replica | PASSIVE
analytics_subcluster | v_verticadb_node0004 | segment0001 | PASSIVE
analytics_subcluster | v_verticadb_node0004 | segment0003 | PASSIVE
analytics_subcluster | v_verticadb_node0005 | replica | PASSIVE
analytics_subcluster | v_verticadb_node0005 | segment0001 | PASSIVE
analytics_subcluster | v_verticadb_node0005 | segment0002 | PASSIVE
analytics_subcluster | v_verticadb_node0006 | replica | PASSIVE
analytics_subcluster | v_verticadb_node0006 | segment0002 | PASSIVE
analytics_subcluster | v_verticadb_node0006 | segment0003 | PASSIVE
default_subcluster | v_verticadb_node0001 | replica | ACTIVE
default_subcluster | v_verticadb_node0001 | segment0001 | ACTIVE
default_subcluster | v_verticadb_node0001 | segment0003 | ACTIVE
default_subcluster | v_verticadb_node0002 | replica | ACTIVE
default_subcluster | v_verticadb_node0002 | segment0001 | ACTIVE
default_subcluster | v_verticadb_node0002 | segment0002 | ACTIVE
default_subcluster | v_verticadb_node0003 | replica | ACTIVE
default_subcluster | v_verticadb_node0003 | segment0002 | ACTIVE
default_subcluster | v_verticadb_node0003 | segment0003 | ACTIVE
(18 rows)
In the previous example, the state of the subscription for the new nodes is PASSIVE. This state means they are warming their depots to prepare for queries. This state eventually becomes ACTIVE, when the nodes have filled their depot with recently-used data. The amount of time this process takes depends on the size of the depot and how frequently data is being evicted from the depot.
Removing Nodes
To remove one or more nodes, use admintools's db_remove_node tool:
$ adminTools -t db_remove_node -p 'password' -d verticadb -s 10.11.12.117
connecting to 10.11.12.10
Waiting for rebalance shards. We will wait for at most 36000 seconds.
Rebalance shards polling iteration number [0], started at [14:56:41], time out at [00:56:41]
Attempting to drop node v_verticadb_node0004 ( 10.11.12.117 )
Shutting down node v_verticadb_node0004
Sending node shutdown command to '['v_verticadb_node0004', '10.11.12.117', '/vertica/data', '/vertica/data']'
Deleting catalog and data directories
Update admintools metadata for v_verticadb_node0004
Eon mode detected. The node v_verticadb_node0004 has been removed from host 10.11.12.117. To remove the
node metadata completely, please clean up the files corresponding to this node, at the communal
location: s3://eonbucket/metadata/verticadb/nodes/v_verticadb_node0004
Reload spread configuration
Replicating configuration to all nodes
Checking database state
Node Status: v_verticadb_node0001: (UP) v_verticadb_node0002: (UP) v_verticadb_node0003: (UP)
Communal storage detected: syncing catalog
When you remove one or more nodes from a subcluster, Vertica automatically rebalances shards in the subcluster. You do not need to manually rebalance shards after removing nodes.
Moving Nodes Between Subclusters
To move a node from one subcluster to another:
- Remove the node or nodes from the subcluster it is currently a part of.
- Add the node to the subcluster you want to move it to.