# Aql client overwhelmed by "WARN AEROSPIKE\_ERR\_TIMEOUT"

**URL:** <https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787>\
**Category:** AQL\
**Created:** [January 23, 2017, 6:20am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787 "2017-01-23T06:20:15Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![billbargens](https://avatars.discourse-cdn.com/v4/letter/b/a5b964/32.png) [@billbargens](https://discuss.aerospike.com/u/billbargens)\
**Post date:** [January 23, 2017, 6:20am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/1 "2017-01-23T06:20:15Z")

</div>

I have a aerospike cluster of two nodes marked as nodeA and nodeB. The cluster have been running normally for a long time.

But Recently when I connected to the server node using aql, the client was overwhelmed by the warning messages:“WARN AEROSPIKE\_ERR\_TIMEOUT”.

 ![](https://us1.discourse-cdn.com/flex019/uploads/aerospike/original/1X/09d95e86ca9c362692af2e32b73ea81f86414913.png)

I run some commands to check the issue, and the result is as follows:

> **asinfo -v service -h nodeA -p 6000** nodeA-ip1:6000;nodeA-ip2:6000

> **asinfo -v service -h nodeB -p 6000** nodeB-ip1:6000;nodeB-ip2:6000; **172.17.42.1:6000**

**The “172.17.42.1:6000” is a unknow ip that does not belongs to nodeB.**

> **asadm -e “asinfo -v services” -p 6000**

> nodeA returned: nodeB-ip1;nodeB-ip2;172.17.42.1:6000

> nodeB returned: nodeA-ip1;nodeA-ip2

> 172 (172.17.42.1) returned: Invalid command or Could not connect to node 172.17.42.1

So how can I get rid of the warning messages?

---

<div class="post-metadata">

**Author:** ![billbargens](https://avatars.discourse-cdn.com/v4/letter/b/a5b964/32.png) [@billbargens](https://discuss.aerospike.com/u/billbargens)\
**Post date:** [January 25, 2017, 8:23am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/2 "2017-01-25T08:23:21Z")

</div>

Any help will be appreciated!

---

<div class="post-metadata">

**Author:** ![billbargens](https://avatars.discourse-cdn.com/v4/letter/b/a5b964/32.png) [@billbargens](https://discuss.aerospike.com/u/billbargens)\
**Post date:** [February 6, 2017, 3:38am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/3 "2017-02-06T03:38:00Z")

</div>

@rbotzer Can you give me some advice?

---

<div class="post-metadata">

**Author:** ![billbargens](https://avatars.discourse-cdn.com/v4/letter/b/a5b964/32.png) [@billbargens](https://discuss.aerospike.com/u/billbargens)\
**Post date:** [February 6, 2017, 3:43am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/4 "2017-02-06T03:43:57Z")

</div>

@system Can anyone give me some advice?

---

<div class="post-metadata">

**Author:** ![kporter](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/kporter/32/515_2.png) [@kporter](https://discuss.aerospike.com/u/kporter)\
**Post date:** [February 6, 2017, 3:57am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/5 "2017-02-06T03:57:33Z")

</div>

If that isn’t a recognized address could it be a rogue node?

---

<div class="post-metadata">

**Author:** ![pgupta](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/pgupta/32/2351_2.png) [@pgupta](https://discuss.aerospike.com/u/pgupta)\
**Post date:** [February 6, 2017, 5:46am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/6 "2017-02-06T05:46:50Z")

</div>

Can you share the exact config files of both nodes? you can mask the ip address if you want to with “nodeA” , “nodeB” etc.

Or you may try to run on nodeA:

> asinfo -v “services-alumni-reset”

and see if that clears it up.

---

<div class="post-metadata">

**Author:** ![billbargens](https://avatars.discourse-cdn.com/v4/letter/b/a5b964/32.png) [@billbargens](https://discuss.aerospike.com/u/billbargens)\
**Post date:** [February 6, 2017, 6:38am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/7 "2017-02-06T06:38:40Z")

</div>

My server version is 3.5.8

The command failed with a exception:

> asinfo -v “services-alumni-reset” -h nodeA request to nodeA returned error

---

<div class="post-metadata">

**Author:** ![billbargens](https://avatars.discourse-cdn.com/v4/letter/b/a5b964/32.png) [@billbargens](https://discuss.aerospike.com/u/billbargens)\
**Post date:** [February 6, 2017, 6:47am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/9 "2017-02-06T06:47:40Z")

</div>

The config files is as fellows:

```auto
 service {
 user root
 group root
  paxos-single-replica-limit 1 # Number of nodes where the replica count is automatically reduced to 1.
  pidfile /disk1/jijw/aerospike/item/var/run/aerospike.pid
  transaction-queues 32
  transaction-threads-per-queue 32
  service-threads 32
  proto-fd-max 15000
  work-directory /disk1/jijw/aerospike/item/var
}

logging {
  # Log file must be an absolute path.
  file /disk1/jijw/aerospike/item/logs/aerospike.log {
    context any info
  }
}

mod-lua {
  system-path /disk1/jijw/aerospike/item/share/udf/lua
  user-path /disk1/jijw/aerospike/item/var/udf/lua
}

network {
  service {
    address any
    port 6000
    reuse-address
  }

  heartbeat {
    mode multicast
    address ****** (masked)
    port 9921
    interval 150
    timeout 10
  }

  fabric {
    port 6001
  }

  info {
    port 6003
  }
}

namespace item {
  single-bin false
  replication-factor 2
  memory-size 20G
  default-ttl 0 # 30 days, use 0 to never expire/evict.
  high-water-memory-pct 85
  high-water-disk-pct 85
  stop-writes-pct 90
  write-commit-level-override all
  storage-engine device {
    device /dev/sdc1 # raw device.# device /dev/<device> # (optional) another raw device.
    write-block-size 1M
    data-in-memory false
    cold-start-empty true
  }
}

```

---

<div class="post-metadata">

**Author:** ![pgupta](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/pgupta/32/2351_2.png) [@pgupta](https://discuss.aerospike.com/u/pgupta)\
**Post date:** [February 6, 2017, 7:26am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/10 "2017-02-06T07:26:49Z")

</div>

Since your server is listening on port 6000 instead of default 3000, add -p 6000 to your asinfo command.

Unrelated, cold-start-empty true leaves you vulnerable to losing all your data should the entire cluster restart after a cluster wide fault. Hope you understand the implications of having that in your config file. What you are saying that always ignore the data in the persistent storage medium when booting this node up. This is generally not recommended.

---

<div class="post-metadata">

**Author:** ![billbargens](https://avatars.discourse-cdn.com/v4/letter/b/a5b964/32.png) [@billbargens](https://discuss.aerospike.com/u/billbargens)\
**Post date:** [February 6, 2017, 7:44am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/11 "2017-02-06T07:44:33Z")

</div>

Still same:

> asinfo -v “services-alumni-reset” -h nodeA -p 6000 request to nodeA : 6000 returned error

My server version is 3.5.8. Does it support this command?

Also thanks for your advice about the cold-start-empty configuration!

---

<div class="post-metadata">

**Author:** ![pgupta](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/pgupta/32/2351_2.png) [@pgupta](https://discuss.aerospike.com/u/pgupta)\
**Post date:** [February 6, 2017, 7:55am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/12 "2017-02-06T07:55:29Z")

</div>

At ver - 3.9.1, dun was deprecated and services-alumni-reset introduced. so yes, 3.5.8, will not work

asinfo -v ‘dun:nodes=BB936F106CA0568’ where BB… is the nodeid that you want to remove.

what does asadm\>info show? do you see the rogue node id?

[https://discuss.aerospike.com/t/faq-how-can-a-node-be-removed-from-a-cluster-in-aerospike-3-9-1-and-higher/3333](https://discuss.aerospike.com/t/faq-how-can-a-node-be-removed-from-a-cluster-in-aerospike-3-9-1-and-higher/3333)

---

<div class="post-metadata">

**Author:** ![billbargens](https://avatars.discourse-cdn.com/v4/letter/b/a5b964/32.png) [@billbargens](https://discuss.aerospike.com/u/billbargens)\
**Post date:** [February 6, 2017, 8:49am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/13 "2017-02-06T08:49:15Z")

</div>

The “asadm\>info” result is as fellows:

 ![](https://us1.discourse-cdn.com/flex019/uploads/aerospike/original/1X/0907b4e09eb17b6adca4b7154582da5690a90c1e.png)

---

<div class="post-metadata">

**Author:** ![kporter](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/kporter/32/515_2.png) [@kporter](https://discuss.aerospike.com/u/kporter)\
**Post date:** [February 6, 2017, 7:21pm UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/14 "2017-02-06T19:21:02Z")

</div>

In your configuration, under network.service, set [access-address](http://www.aerospike.com/docs/reference/configuration#access-address) to the appropriate client reachable address. This configuration is static so you will need to restart each node after configuring.

---

<div class="post-metadata">

**Author:** ![billbargens](https://avatars.discourse-cdn.com/v4/letter/b/a5b964/32.png) [@billbargens](https://discuss.aerospike.com/u/billbargens)\
**Post date:** [February 7, 2017, 3:22am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/15 "2017-02-07T03:22:17Z")

</div>

But the request using java client is normal!

---

<div class="post-metadata">

**Author:** ![billbargens](https://avatars.discourse-cdn.com/v4/letter/b/a5b964/32.png) [@billbargens](https://discuss.aerospike.com/u/billbargens)\
**Post date:** [February 7, 2017, 3:29am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/16 "2017-02-07T03:29:10Z")

</div>

The aerospike cluster seems to be running normally except for the endless aql exception!

---

<div class="post-metadata">

**Author:** ![kporter](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/kporter/32/515_2.png) [@kporter](https://discuss.aerospike.com/u/kporter)\
**Post date:** [February 7, 2017, 3:47am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/17 "2017-02-07T03:47:16Z")

</div>

Run:

```auto
asadm -e "asinfo -v service" -p 6000

```

This will show the services each node is advertising.

---

<div class="post-metadata">

**Author:** ![pgupta](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/pgupta/32/2351_2.png) [@pgupta](https://discuss.aerospike.com/u/pgupta)\
**Post date:** [February 7, 2017, 3:50am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/18 "2017-02-07T03:50:14Z")

</div>

I think your cluster keeps looking for this non existent node at 172:17:42:1:6000. Try: asinfo -v ‘tip-clear:host-port-list=172.17.42.1:6000’ -h nodeA -p 6000

Then see if asadm\>info shows only the two good nodes. Also, good idea to do an asbackup of your data before trying anything exotic!

Once you have backup, you can try: asinfo -v ‘dun:nodes=0’ -h nodeA -p 6000 because the nodeid seems to be 0 for this non-existent node.

---

<div class="post-metadata">

**Author:** ![billbargens](https://avatars.discourse-cdn.com/v4/letter/b/a5b964/32.png) [@billbargens](https://discuss.aerospike.com/u/billbargens)\
**Post date:** [February 7, 2017, 5:46am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/19 "2017-02-07T05:46:29Z")

</div>

I run some commands to check the issue, and the result is as follows:

> asinfo -v service -h nodeA -p 6000 nodeA-ip1:6000;nodeA-ip2:6000

> asinfo -v service -h nodeB -p 6000 nodeB-ip1:6000;nodeB-ip2:6000; **172.17.42.1:6000**

The “172.17.42.1:6000” is a unknow ip that does not belongs to nodeB.

> asadm -e “asinfo -v services” -p 6000

> **nodeA returned:** nodeB-ip1;nodeB-ip2;172.17.42.1:6000

> **nodeB returned:** nodeA-ip1;nodeA-ip2

> **172 (172.17.42.1) returned:** Invalid command or Could not connect to node 172.17.42.1

Does this mean the address that nodeB advertised to the cluster was **nodeB-ip1** , **nodeB-ip2** and **172.17.42.1:6000**? But The **“172.17.42.1:6000”** is a unknow ip that does not belongs to nodeB.

---

<div class="post-metadata">

**Author:** ![billbargens](https://avatars.discourse-cdn.com/v4/letter/b/a5b964/32.png) [@billbargens](https://discuss.aerospike.com/u/billbargens)\
**Post date:** [February 7, 2017, 6:09am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/20 "2017-02-07T06:09:53Z")

</div>

I try your methods, but does not work.

As from the official doc: [Info Command Reference | Aerospike Documentation](http://www.aerospike.com/docs/reference/info#service), the command “asinfo -v service -h nodeB -p 6000” will return a list of IP that nodeB advitesd to other cluster nodes.

> asinfo -v service -h nodeB -p 6000 nodeB-ip1:6000;nodeB-ip2:6000;172.17.42.1:6000

It seems that nodeB advertised a ip 172.17.42.1 that does not belong to it to the cluster. So what we need to do is to get rid of that ip. Is it right?

---

<div class="post-metadata">

**Author:** ![pgupta](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/pgupta/32/2351_2.png) [@pgupta](https://discuss.aerospike.com/u/pgupta)\
**Post date:** [February 7, 2017, 6:12am UTC](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787/21 "2017-02-07T06:12:35Z")

</div>

Some other non aerospike process also listening at this port 6000 on nodeB? On node B, can you try using netstat and see what processes are using port 6000?

[Next page](https://discuss.aerospike.com/t/aql-client-overwhelmed-by-warn-aerospike-err-timeout/3787.md?page=2)
