# Invalid\_Node\_Error (Code -3) after some writes

**URL:** https://discuss.aerospike.com/t/invalid-node-error-code-3-after-some-writes/6037
**Category:** Java Client
**Created:** [March 19, 2019, 6:38am UTC](https://discuss.aerospike.com/t/invalid-node-error-code-3-after-some-writes/6037 "2019-03-19T06:38:28Z")
**Posts on this page:** 10
**Page:** 1

<div class="post-metadata">

### Author: ![Pankaj7003](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/pankaj7003/32/1083_2.png) [@Pankaj7003](https://discuss.aerospike.com/u/Pankaj7003)
#### Post date: [March 19, 2019, 6:38am UTC](https://discuss.aerospike.com/t/invalid-node-error-code-3-after-some-writes/6037/1 "2019-03-19T06:38:28Z")

</div>

We’ve a 4 node cluster Aerospike (build 3.12.0 community addition). Our cluster is working fine for read/writes except when we try to write at a high rate through java client via spark job. While executing the job the writes are happening at 500 TPS on all nodes as seen in AMC. But either the job fails after sometime or even if the job passes the second job fails within few seconds with “com.aerospike.client.AerospikeException$InvalidNode: Error Code -3: Invalid node”. Subsequent jobs have the same error for quite sometime (some hours) before next job can write any data. At the time of failure of job cluster seems healthy on all checked metrics, viz. client connection, open connection, pending IO tasks etc.

Following is the config for

```auto
namespace {name_space_name} {
        replication-factor 1
        memory-size 50G
		default-ttl 10D
		storage-engine device {
			file /storage/aerospike/{file_name}.dat
			filesize 350G
			data-in-memory true
        }
}

```

---

<div class="post-metadata">

### Author: ![kporter](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/kporter/32/515_2.png) [@kporter](https://discuss.aerospike.com/u/kporter)
#### Post date: [March 19, 2019, 7:26am UTC](https://discuss.aerospike.com/t/invalid-node-error-code-3-after-some-writes/6037/2 "2019-03-19T07:26:31Z")

</div>

The rest of your config would be more helpful. Likely the old paxos or heartbeat implementations hit an issue. In your version, you can upgrade the heartbeat protocol to v3, which may help. But I suggest upgrading to 3.13 instead. 3.13 reworks much of the distributed system. Also since you must upgrade through 3.13 from prior versions, it has had large extension to the period where we backport bug fixes. You can find instructions here: [https://www.aerospike.com/docs/operations/upgrade/cluster\_to\_3\_13/](https://www.aerospike.com/docs/operations/upgrade/cluster_to_3_13/).

Additional information about 3.13 can be found here: [What’s New in Aerospike 3.13 and 3.14? | Aerospike](https://www.aerospike.com/blog/whats-new-aerospike-3-13-3-14/).

---

<div class="post-metadata">

### Author: ![Pankaj7003](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/pankaj7003/32/1083_2.png) [@Pankaj7003](https://discuss.aerospike.com/u/Pankaj7003)
#### Post date: [March 19, 2019, 7:39am UTC](https://discuss.aerospike.com/t/invalid-node-error-code-3-after-some-writes/6037/3 "2019-03-19T07:39:07Z")

</div>

Thanks for speedy reply @kporter. Will upgrade the heartbeat protocol and try out first as cluster upgrade is comparatively bigger task. Meanwhile below is rest of the config if helpful:

```auto
network {
        service {
                address any
                port 3000
        }

        heartbeat {
                mode mesh
                port 3002

                mesh-seed-address-port {ip1} 3002
                mesh-seed-address-port {ip2} 3002
                mesh-seed-address-port {ip3} 3002
                mesh-seed-address-port {ip4} 3002

                interval 150
                timeout 20
        }

        fabric {
                port 3001
        }

        info {
                port 3003
        }
}

service {

        user {username}
        group {group_name}

        nsup-period 100
        paxos-single-replica-limit 1
        service-threads 20
        transaction-queues 20
        transaction-threads-per-queue 3
        transaction-pending-limit 15000
        proto-fd-max 50000
        migrate-threads 1
        pidfile /var/run/aerospike/asd.pid
}

```

---

<div class="post-metadata">

### Author: ![Pankaj7003](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/pankaj7003/32/1083_2.png) [@Pankaj7003](https://discuss.aerospike.com/u/Pankaj7003)
#### Post date: [March 19, 2019, 12:40pm UTC](https://discuss.aerospike.com/t/invalid-node-error-code-3-after-some-writes/6037/4 "2019-03-19T12:40:04Z")

</div>

> [@kporter](#):
>
> In your version, you can upgrade the heartbeat protocol to v3, which may help

Unfortunately it didn’t help out and its still same. Also, what I missed in last post was that if we restart the Aerospike nodes, cluster again starts to accept bulk writes. Please let me know if upgrading would be the only viable option.

---

<div class="post-metadata">

### Author: ![meher](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/meher/32/1644_2.png) [@meher](https://discuss.aerospike.com/u/meher)
#### Post date: [March 24, 2019, 6:22pm UTC](https://discuss.aerospike.com/t/invalid-node-error-code-3-after-some-writes/6037/5 "2019-03-24T18:22:15Z")

</div>

This error -3 invalid node could be caused by attempting connections without having the cluster object instantiated (no nodes have been discovered in the cluster). Also, are you using the latest Java Client?

---

<div class="post-metadata">

### Author: ![Pankaj7003](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/pankaj7003/32/1083_2.png) [@Pankaj7003](https://discuss.aerospike.com/u/Pankaj7003)
#### Post date: [April 2, 2019, 6:09am UTC](https://discuss.aerospike.com/t/invalid-node-error-code-3-after-some-writes/6037/6 "2019-04-02T06:09:44Z")

</div>

> attempting connections without having the cluster object instantiated

In that case there shouldn’t be any updates going at all, but writes do go for sometime and then fails. BTW, when would this scenario happen ?

> Also, are you using the latest Java Client?

We’re using “ **3.3.2** ” Java Client. Will check if there are no breaking changes in the latest client.

---

<div class="post-metadata">

### Author: ![meher](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/meher/32/1644_2.png) [@meher](https://discuss.aerospike.com/u/meher)
#### Post date: [April 3, 2019, 1:17am UTC](https://discuss.aerospike.com/t/invalid-node-error-code-3-after-some-writes/6037/7 "2019-04-03T01:17:30Z")

</div>

I am not an expert on the client coding best practices, but I guess it could happen on how the cluster object is initialized / re-initialized or when recovering from a short lived network outage?

I would definitely suggest trying the latest client. There would be some changes to apply but I don’t think it is much. The release notes would have links to the relevant details.

---

<div class="post-metadata">

### Author: ![Nipun\_Jain](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/nipun_jain/32/1243_2.png) [@Nipun\_Jain](https://discuss.aerospike.com/u/Nipun_Jain)
#### Post date: [August 17, 2019, 5:27am UTC](https://discuss.aerospike.com/t/invalid-node-error-code-3-after-some-writes/6037/8 "2019-08-17T05:27:32Z")

</div>

@Pankaj7003 How’d you resolve this issue? I am facing the same issue. Using Community Edition - 4.0.19

---

<div class="post-metadata">

### Author: ![Brian](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/brian/32/2211_2.png) [@Brian](https://discuss.aerospike.com/u/Brian)
#### Post date: [August 19, 2019, 4:59pm UTC](https://discuss.aerospike.com/t/invalid-node-error-code-3-after-some-writes/6037/9 "2019-08-19T16:59:35Z")

</div>

I recommend upgrading to the latest java client.

---

<div class="post-metadata">

### Author: ![system](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/system/32/2274_2.png) [@system](https://discuss.aerospike.com/u/system)
#### Post date: [August 25, 2019, 5:10pm UTC](https://discuss.aerospike.com/t/invalid-node-error-code-3-after-some-writes/6037/10 "2019-08-25T17:10:13Z")

</div>

This topic was automatically closed 6 days after the last reply. New replies are no longer allowed.
