# Write should continue if one of the node fails

**URL:** <https://discuss.aerospike.com/t/write-should-continue-if-one-of-the-node-fails/5757>\
**Category:** How Aerospike Works\
**Created:** [December 7, 2018, 1:30pm UTC](https://discuss.aerospike.com/t/write-should-continue-if-one-of-the-node-fails/5757 "2018-12-07T13:30:09Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mahaveer\_Darade](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/mahaveer_darade/32/1060_2.png) [@Mahaveer\_Darade](https://discuss.aerospike.com/u/Mahaveer_Darade)\
**Post date:** [December 7, 2018, 1:30pm UTC](https://discuss.aerospike.com/t/write-should-continue-if-one-of-the-node-fails/5757/1 "2018-12-07T13:30:09Z")

</div>

Hi All,

I am using latest client & server (community edition). If one of the node from cluster (having replication factor as 2) fails, I want write op to continue & eventually complete.

I’ve tried setting below policies.

cfg.policies.read.replica = AS\_POLICY\_REPLICA\_ANY; cfg.policies.write.replica = AS\_POLICY\_REPLICA\_ANY; cfg.policies.write.commit\_level = AS\_POLICY\_COMMIT\_LEVEL\_MASTER;

cfg.policies.write.base.total\_timeout = 5000;// 5secs cfg.policies.write.base.max\_retries = 1; cfg.policies.write.base.sleep\_between\_retries = 300; // 300ms

But I don’t see write operation getting completed. Its getting stuck forever. Can anyone please help with some pointers?

---

<div class="post-metadata">

**Author:** ![kporter](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/kporter/32/515_2.png) [@kporter](https://discuss.aerospike.com/u/kporter)\
**Post date:** [December 7, 2018, 6:50pm UTC](https://discuss.aerospike.com/t/write-should-continue-if-one-of-the-node-fails/5757/2 "2018-12-07T18:50:29Z")

</div>

All modes should achieve this result, how are you determining that the write isn’t completing?

What is meant by ‘stuck forever’?

---

<div class="post-metadata">

**Author:** ![Mahaveer\_Darade](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/mahaveer_darade/32/1060_2.png) [@Mahaveer\_Darade](https://discuss.aerospike.com/u/Mahaveer_Darade)\
**Post date:** [December 13, 2018, 12:37pm UTC](https://discuss.aerospike.com/t/write-should-continue-if-one-of-the-node-fails/5757/4 "2018-12-13T12:37:54Z")

</div>

Hi Kevin,   
I’ve modified async\_get example to write 30000 records.   
It works fine as I see 60000 records into DB (considering replication factor of 2).

```auto
Admin> show statistics sets 
~~~~~~~~~~~~~~~~~~~~~~~~DIRTY eg-set Set Statistics (2018-12-13 11:51:33 UTC)~~~~~~~~~~~~~~~~~~~~~~~~
NODE : 192.168.4.10:3000 192.168.4.164:3000 192.168.6.192:3000 192.168.7.186:3000   
disable-eviction : false false false false                
memory_data_bytes: 0 0 0 0                    
ns : DIRTY DIRTY DIRTY DIRTY                
objects : 14720 15161 14689 15430                
set : eg-set eg-set eg-set eg-set               
set-enable-xdr : use-default use-default use-default use-default          
stop-writes-count: 0 0 0 0                    
tombstones : 0 0 0 0                    
truncate_lut : 0 0 0 0                    

Admin> 

```

Now when I bring down 192.168.4.10 (while writes still happening) I see less records than 60000.

```auto
Admin> show statistics sets 
~~~~~~~~~~~~~~DIRTY eg-set Set Statistics (2018-12-13 12:27:52 UTC)~~~~~~~~~~~~~~
NODE : 192.168.4.164:3000 192.168.6.192:3000 192.168.7.186:3000   
disable-eviction : false false false                
memory_data_bytes: 0 0 0                    
ns : DIRTY DIRTY DIRTY                
objects : 19821 20093 20006                
set : eg-set eg-set eg-set               
set-enable-xdr : use-default use-default use-default          
stop-writes-count: 0 0 0                    
tombstones : 0 0 0                    
truncate_lut : 0 0 0                    

Admin> 

```

I get below logs from aerospike c client which detects node removal from cluster.

```auto
write succeeded 5000
[src/main/aerospike/as_cluster.c:547][as_cluster_tend] 2 - Node BB91ABB97565000 refresh failed: AEROSPIKE_ERR_CONNECTION Bad file descriptor
[src/main/aerospike/as_cluster.c:547][as_cluster_tend] 2 - Node BB91ABB97565000 refresh failed: AEROSPIKE_ERR_CONNECTION Socket write error: 111, 192.168.4.10:3000, 35954
[src/main/aerospike/as_cluster.c:547][as_cluster_tend] 2 - Node BB91ABB97565000 refresh failed: AEROSPIKE_ERR_CONNECTION Socket write error: 111, 192.168.4.10:3000, 35956
[src/main/aerospike/as_cluster.c:362][as_cluster_remove_nodes_copy] 2 - Remove node BB91ABB97565000 192.168.4.10:3000
count:5001, ev-loop:0, pending-cmds:0

```

```auto
write succeeded 15000
[src/main/aerospike/as_node.c:158][as_node_destroy] 2 - as_node_destroy start for node 192.168.4.10:3000
[src/main/aerospike/as_node.c:205][as_node_destroy] 2 - as_node_destroy end for node 192.168.4.10:3000
count:15001, ev-loop:0, pending-cmds:0

```

End goal is to get 60000 records into DB even if one of the node fails while writes happening.   
Is there any configuration policy for aerospike c client which can retry write op to different node?

---

<div class="post-metadata">

**Author:** ![Albot](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/albot/32/2076_2.png) [@Albot](https://discuss.aerospike.com/u/Albot)\
**Post date:** [December 15, 2018, 4:11am UTC](https://discuss.aerospike.com/t/write-should-continue-if-one-of-the-node-fails/5757/5 "2018-12-15T04:11:31Z")

</div>

You could quiesce the node before taking it down, that way it would be graceful. Otherwise, check to see if the c client supports retries and/or build your own retry logic.

---

<div class="post-metadata">

**Author:** ![kporter](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/kporter/32/515_2.png) [@kporter](https://discuss.aerospike.com/u/kporter)\
**Post date:** [December 17, 2018, 7:34pm UTC](https://discuss.aerospike.com/t/write-should-continue-if-one-of-the-node-fails/5757/6 "2018-12-17T19:34:10Z")

</div>

When you take down a node, you will be missing a copy of some of your records. If you wait for migrations to complete, your unique-data \* replication-factor objects will be restored.

And when you take down a node, there will be connection errors. You can eliminate these errors in planned events using the [`quiesce`](https://newwww.aerospike.com/docs/operations/manage/cluster_mng/quiescing_node/index.html) feature @Albot mentioned.

For unplanned events, the [`AS_POLICY_REPLICA_SEQUENCE`](https://www.aerospike.com/apidocs/c/db/d65/group __client__ policies.html#gabce1fb468ee9cbfe54b7ab834cec79ab) policy will try the master partition first and next, if retries are enabled, it will try the next prole partition in sequence. The purpose of the `REPLICA_SEQUENCE` policy is to minimize these exceptions on unplanned events since the next prole is likely to become master if the master did drop. (I do not know of good reason to ever use `REPLICA_ANY`.)

You could also implement a custom retry mechanism on the client. I would suggest not retrying errors when [`in_doubt == false`](https://www.aerospike.com/apidocs/c/d8/d36/structas__error.html#aab83b7d1f4bea75a2ad98d1ec81f93eb) because this means that the current transaction cannot succeed.
