# Numerous timeout exceptions with PHP Client

**URL:** <https://discuss.aerospike.com/t/numerous-timeout-exceptions-with-php-client/3863>\
**Category:** PHP Client Library\
**Created:** [February 16, 2017, 4:46pm UTC](https://discuss.aerospike.com/t/numerous-timeout-exceptions-with-php-client/3863 "2017-02-16T16:46:45Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![Crags](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/crags/32/309_2.png) [@Crags](https://discuss.aerospike.com/u/Crags)\
**Post date:** [February 16, 2017, 4:46pm UTC](https://discuss.aerospike.com/t/numerous-timeout-exceptions-with-php-client/3863/1 "2017-02-16T16:46:45Z")

</div>

Hi all

I’m currently getting quite a few timeout exceptions even though we have a timeout set to 5 seconds, and we’re simply grabbing 1 key. Admittedly we have around 20k connections but I would assume Aerospike could handle this no problem? Here’s a snippet:

```auto
Thu, 16 Feb 2017 16:30:20 +0000	FATAL	Error while getting data from Aerospike server: Key [{"ns":"adspruce","set":"test_sites","key":"3239-1"}] : Error - [9]: Timeout: timeout=5000 iterations=1 failedNodes=0 failedConns=0	
Thu, 16 Feb 2017 16:30:23 +0000	FATAL	Error while getting data from Aerospike server: Key [{"ns":"adspruce","set":"test_sites","key":"3239-1"}] : Error - [9]: Timeout: timeout=5000 iterations=1 failedNodes=0 failedConns=0	
Thu, 16 Feb 2017 16:31:30 +0000	FATAL	Error while getting data from Aerospike server: Key [{"ns":"adspruce","set":"test_sites","key":"3086-2"}] : Error - [9]: Timeout: timeout=5000 iterations=1 failedNodes=0 failedConns=0	
Thu, 16 Feb 2017 16:32:21 +0000	FATAL	Error while getting data from Aerospike server: Key [{"ns":"adspruce","set":"test_sites","key":"3239-1"}] : Error - [9]: Timeout: timeout=5000 iterations=1 failedNodes=0 failedConns=0	
Thu, 16 Feb 2017 16:33:05 +0000	FATAL	Error while getting data from Aerospike server: Key [{"ns":"adspruce","set":"test_sites","key":"3239-1"}] : Error - [9]: Timeout: timeout=5000 iterations=1 failedNodes=0 failedConns=0	
Thu, 16 Feb 2017 16:33:34 +0000	FATAL	Error while getting data from Aerospike server: Key [{"ns":"adspruce","set":"test_sites","key":"3239-1"}] : Error - [9]: Timeout: timeout=5000 iterations=1 failedNodes=0 failedConns=0	
Thu, 16 Feb 2017 16:35:37 +0000	FATAL	Error while getting data from Aerospike server: Key [{"ns":"adspruce","set":"test_sites","key":"3088-2"}] : Error - [9]: Timeout: timeout=5000 iterations=1 failedNodes=0 failedConns=0	
Thu, 16 Feb 2017 16:35:37 +0000	FATAL	Error while getting data from Aerospike server: Key [{"ns":"temp01","set":"device_data","key":"7e803f93ee144b46f70c54bb9a00664a35604476"}] : Error - [9]: Timeout: timeout=5000 iterations=1 failedNodes=0 failedConns=0	
Thu, 16 Feb 2017 16:35:58 +0000	FATAL	Error while getting data from Aerospike server: Key [{"ns":"adspruce","set":"test_sites","key":"3086-2"}] : Error - [9]: Timeout: timeout=5000 iterations=1 failedNodes=0 failedConns=0	
Thu, 16 Feb 2017 16:36:31 +0000	FATAL	Error while getting data from Aerospike server: Key [{"ns":"adspruce","set":"test_sites","key":"3088-2"}] : Error - [9]: Timeout: timeout=5000 iterations=1 failedNodes=0 failedConns=0	
Thu, 16 Feb 2017 16:39:01 +0000	FATAL	Error while getting data from Aerospike server: Key [{"ns":"adspruce","set":"test_sites","key":"2915-1"}] : Error - [9]: Timeout: timeout=5000 iterations=1 failedNodes=0 failedConns=0

```

Any idea what could be causing these timeout issues, and whether there is anything I can change in config or benchmarks I can run to see if it is client related or whether I should just be increasing the timeout - already 5 seconds seems like long enough

Many thanks

Craig

---

<div class="post-metadata">

**Author:** ![rbotzer](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/rbotzer/32/2114_2.png) [@rbotzer](https://discuss.aerospike.com/u/rbotzer)\
**Post date:** [February 16, 2017, 4:59pm UTC](https://discuss.aerospike.com/t/numerous-timeout-exceptions-with-php-client/3863/2 "2017-02-16T16:59:24Z")

</div>

What’s going on at the server at the same time? Try to correlate the log for the same time period (and/or post it here). It’s rarely a limitation of Aerospike, it has to do with config or the actual hardware. For example, are you seeing ‘queue too deep’ warnings on the server side?

What hardware is the server running on?

---

<div class="post-metadata">

**Author:** ![Albot](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/albot/32/2076_2.png) [@Albot](https://discuss.aerospike.com/u/Albot)\
**Post date:** [February 18, 2017, 3:10am UTC](https://discuss.aerospike.com/t/numerous-timeout-exceptions-with-php-client/3863/3 "2017-02-18T03:10:59Z")

</div>

What is your current connection count vs proto-fd-max? Also, you say you’re just trying to get 1 key, are you able to connect at all (telnet xyz 3000)?

As robert said, an aerospike.log with the same timeframes would be helpful here… (and asadm statistics/conf wouldnt hurt either)

---

<div class="post-metadata">

**Author:** ![Crags](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/crags/32/309_2.png) [@Crags](https://discuss.aerospike.com/u/Crags)\
**Post date:** [February 28, 2017, 10:12am UTC](https://discuss.aerospike.com/t/numerous-timeout-exceptions-with-php-client/3863/4 "2017-02-28T10:12:42Z")

</div>

Sorry @rbotzer - only just getting back to this now after being on leave…

I’ve had to do some log rotations as it wasn’t working (CentOS 7) but can’t see anything in there so far that you’ve mentioned. There was an issue with `shm` not working in the FPM .ini config but I’ve resolved that now, however we still see quite a few timeouts as shown above.

we have 2 x R430 servers running the cluster running with 220Gb Ram - not sure on cores etc at the minute.

To answer @Albot question, we were running at 15k `proto-fd-max` but we hit and surpassed that now to have it running at 50k - our FPM config is running with 500 children handling 10k connections each - and we have 10 front end app servers. When we hit the connection limit I was unable to telnet in, but when we see timeout issues I can telnet / ssh in fine and the servers are running fine.

One other issue we are seeing, is according to checkMK, aerospike seems to be holding on to TCP connections, but in the `ESTABLISHED` state, not the `WAIT` state like others have seen. I’m not sure whether this is normal behaviour or a symptom of something else?

If you need more data / config / asadm outputs let me know.

Thanks both

---

<div class="post-metadata">

**Author:** ![Crags](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/crags/32/309_2.png) [@Crags](https://discuss.aerospike.com/u/Crags)\
**Post date:** [February 28, 2017, 3:04pm UTC](https://discuss.aerospike.com/t/numerous-timeout-exceptions-with-php-client/3863/5 "2017-02-28T15:04:48Z")

</div>

With regards to the connections, over the last week you can see this steady increase - surely it should flow in peaks and troughs with our traffic?:

![](https://us1.discourse-cdn.com/flex019/uploads/aerospike/original/1X/8774261a0c68af9b114e0132ffb8bb0173c0ae55.png)

---

<div class="post-metadata">

**Author:** ![Albot](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/albot/32/2076_2.png) [@Albot](https://discuss.aerospike.com/u/Albot)\
**Post date:** [March 1, 2017, 3:12am UTC](https://discuss.aerospike.com/t/numerous-timeout-exceptions-with-php-client/3863/6 "2017-03-01T03:12:29Z")

</div>

Are you using a lot of secondary index queries? Does your connection from the app to DB go through a firewall?

---

<div class="post-metadata">

**Author:** ![Crags](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/crags/32/309_2.png) [@Crags](https://discuss.aerospike.com/u/Crags)\
**Post date:** [March 1, 2017, 8:33am UTC](https://discuss.aerospike.com/t/numerous-timeout-exceptions-with-php-client/3863/7 "2017-03-01T08:33:33Z")

</div>

We are doing a lot of secondary index queries yes, we basically look up a number of records based on an indexed boolean field being true - but we’re not talking millions of records, mainly hundreds or a few thousand - nothing I would have thought was very taxing on the system.

At the time of writing, we’re doing somewhere between 1k and 1.5k TPS in the query field according to AMC. Also, checking the connections, it simply seems as if they are opening and closing, but not all of them - there is a definite and slow increase of a ‘baseline’ which makes me think _something_ is not closing them properly, or there is an issue with the client possibly which may be doing it?

Thanks again

---

<div class="post-metadata">

**Author:** ![Albot](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/albot/32/2076_2.png) [@Albot](https://discuss.aerospike.com/u/Albot)\
**Post date:** [March 1, 2017, 5:58pm UTC](https://discuss.aerospike.com/t/numerous-timeout-exceptions-with-php-client/3863/8 "2017-03-01T17:58:48Z")

</div>

If you restart the app, and reap any old connections, do you have issues on initial startup?

---

<div class="post-metadata">

**Author:** ![Crags](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/crags/32/309_2.png) [@Crags](https://discuss.aerospike.com/u/Crags)\
**Post date:** [March 1, 2017, 8:26pm UTC](https://discuss.aerospike.com/t/numerous-timeout-exceptions-with-php-client/3863/9 "2017-03-01T20:26:05Z")

</div>

I have just restarted the entire cluster, and we now have the following:

![](https://us1.discourse-cdn.com/flex019/uploads/aerospike/original/1X/63d661a2a4d655be168d625be30d1f37b30e574e.jpg)

This is a very quiet period for us, so this is to be expected - I will keep an eye on it now and see how it increases and how quickly - it simply seems as if some connections are not being released and these are multiplying over time

---

<div class="post-metadata">

**Author:** ![Crags](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/crags/32/309_2.png) [@Crags](https://discuss.aerospike.com/u/Crags)\
**Post date:** [March 2, 2017, 8:40am UTC](https://discuss.aerospike.com/t/numerous-timeout-exceptions-with-php-client/3863/10 "2017-03-02T08:40:07Z")

</div>

Okay, so just checked now and connections is up to around 1k on both nodes, throughout and traffic still the same as it was 12 hours ago. Just seems to be steadily and slowly increasing.

I’m wondering if some of the fixes in 3.4.14 have not worked?

```auto
Fix persistent connection issue under php-fpm which caused number of threads to increase
Fix shared memory persistence under php-fpm

```

---

<div class="post-metadata">

**Author:** ![Crags](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/crags/32/309_2.png) [@Crags](https://discuss.aerospike.com/u/Crags)\
**Post date:** [March 2, 2017, 2:33pm UTC](https://discuss.aerospike.com/t/numerous-timeout-exceptions-with-php-client/3863/11 "2017-03-02T14:33:01Z")

</div>

you can see here the gradual increase… this is the last 4 hours

![](https://us1.discourse-cdn.com/flex019/uploads/aerospike/original/1X/e19780f134525b535d356d5ddee269d2b96f0b5b.png)

---

<div class="post-metadata">

**Author:** ![Albot](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/albot/32/2076_2.png) [@Albot](https://discuss.aerospike.com/u/Albot)\
**Post date:** [March 3, 2017, 12:33am UTC](https://discuss.aerospike.com/t/numerous-timeout-exceptions-with-php-client/3863/12 "2017-03-03T00:33:24Z")

</div>

There was a recent bug reported for secondary index pipes not being closed correctly that show similar patterns for me. I think they are releasing a fix in the next version 🙂

---

<div class="post-metadata">

**Author:** ![Albot](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/albot/32/2076_2.png) [@Albot](https://discuss.aerospike.com/u/Albot)\
**Post date:** [March 3, 2017, 12:34am UTC](https://discuss.aerospike.com/t/numerous-timeout-exceptions-with-php-client/3863/13 "2017-03-03T00:34:22Z")

</div>

Again though, do you see the issues when retrying to reproduce the issue after restarting the application?

---

<div class="post-metadata">

**Author:** ![Crags](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/crags/32/309_2.png) [@Crags](https://discuss.aerospike.com/u/Crags)\
**Post date:** [March 3, 2017, 9:39am UTC](https://discuss.aerospike.com/t/numerous-timeout-exceptions-with-php-client/3863/14 "2017-03-03T09:39:50Z")

</div>

Basically, I restarted at around 8pm on the 2nd March, and the connections have been steadily increasing ever since over and above any peak periods.

This would make sense in terms of your bug report, as around 50% of the queries we are running are secondary index queries. Any ETA on the release of the fix?

---

<div class="post-metadata">

**Author:** ![Albot](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/albot/32/2076_2.png) [@Albot](https://discuss.aerospike.com/u/Albot)\
**Post date:** [March 6, 2017, 5:44am UTC](https://discuss.aerospike.com/t/numerous-timeout-exceptions-with-php-client/3863/15 "2017-03-06T05:44:01Z")

</div>

the connection issue may be unrelated to your issue. Did your app have issues after being restarted??

---

<div class="post-metadata">

**Author:** ![Crags](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/crags/32/309_2.png) [@Crags](https://discuss.aerospike.com/u/Crags)\
**Post date:** [March 6, 2017, 8:51am UTC](https://discuss.aerospike.com/t/numerous-timeout-exceptions-with-php-client/3863/16 "2017-03-06T08:51:47Z")

</div>

Yeah, we’re still getting timeout issues with requests
