# Writes failing to aerospike cluster

**URL:** <https://discuss.aerospike.com/t/writes-failing-to-aerospike-cluster/4009>\
**Category:** General Discussion\
**Created:** [March 31, 2017, 12:22am UTC](https://discuss.aerospike.com/t/writes-failing-to-aerospike-cluster/4009 "2017-03-31T00:22:30Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![sdas](https://avatars.discourse-cdn.com/v4/letter/s/bc79bd/32.png) [@sdas](https://discuss.aerospike.com/u/sdas)\
**Post date:** [March 31, 2017, 12:22am UTC](https://discuss.aerospike.com/t/writes-failing-to-aerospike-cluster/4009/1 "2017-03-31T00:22:30Z")

</div>

Hi,

I am new to using aerospike. For my use case, I am reading data from EMR and writing to an aerospike cluster of 2 AWS instances which are of m4.2xlarge instance types. One thing that I am noticing is that the successfull TPS is almost half of the total TPS during the write job. I am attaching the snapshot for reference. So if the messages are actually getting dropped, is there any paramter in the Aerospike client that I can use to guard against it.

Thanks,

![](https://us1.discourse-cdn.com/flex019/uploads/aerospike/original/1X/6290c66872a409831e65bfef72c64a33d323a91f.png)

---

<div class="post-metadata">

**Author:** ![kporter](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/kporter/32/515_2.png) [@kporter](https://discuss.aerospike.com/u/kporter)\
**Post date:** [March 31, 2017, 12:47am UTC](https://discuss.aerospike.com/t/writes-failing-to-aerospike-cluster/4009/2 "2017-03-31T00:47:26Z")

</div>

1. Aerospike version?

2. Grep logs for a line similar to:

3. Share namespace stats:

---

<div class="post-metadata">

**Author:** ![sdas](https://avatars.discourse-cdn.com/v4/letter/s/bc79bd/32.png) [@sdas](https://discuss.aerospike.com/u/sdas)\
**Post date:** [March 31, 2017, 1:45am UTC](https://discuss.aerospike.com/t/writes-failing-to-aerospike-cluster/4009/3 "2017-03-31T01:45:40Z")

</div>

Hi,

I am using Aerospike version 3.10.1.1.

I am attaching the screenshot for the logs when the client was doing the writes

 ![](https://us1.discourse-cdn.com/flex019/uploads/aerospike/original/1X/839016c84de31c1e9dc8e8914c095d495d55f69a.png)

For node 2 it was

Mar 31 2017 01:21:46 GMT: INFO (info): (ticker.c:551) {namespace-dev} client: tsvc (0,0) proxy (0,0,0) read (0,0,0,0) write (19893,4454,0) delete (0,0,0,0) udf (0,0,0) lang (0,0,0,0) Mar 31 2017 01:21:56 GMT: INFO (info): (ticker.c:551) {namespace-dev} client: tsvc (0,0) proxy (0,0,0) read (0,0,0,0) write (19893,4454,0) delete (0,0,0,0) udf (0,0,0) lang (0,0,0,0) Mar 31 2017 01:22:06 GMT: INFO (info): (ticker.c:551) {namespace-dev} client: tsvc (0,0) proxy (0,0,0) read (0,0,0,0) write (19893,4454,0) delete (0,0,0,0) udf (0,0,0) lang (0,0,0,0)

Also if you can tell me which parameters from the stat command you are specifically interested in, I can add them to the comment.

Thanks, Soudipta

---

<div class="post-metadata">

**Author:** ![kporter](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/kporter/32/515_2.png) [@kporter](https://discuss.aerospike.com/u/kporter)\
**Post date:** [March 31, 2017, 4:45am UTC](https://discuss.aerospike.com/t/writes-failing-to-aerospike-cluster/4009/4 "2017-03-31T04:45:08Z")

</div>

```auto
asadm -e "show stat namespace for <namespace name> like client write fail"

```

---

<div class="post-metadata">

**Author:** ![sdas](https://avatars.discourse-cdn.com/v4/letter/s/bc79bd/32.png) [@sdas](https://discuss.aerospike.com/u/sdas)\
**Post date:** [March 31, 2017, 7:25am UTC](https://discuss.aerospike.com/t/writes-failing-to-aerospike-cluster/4009/5 "2017-03-31T07:25:24Z")

</div>

Hi Kevin, Thanks for the pointer. I am attaching the screen shot of the stat command

 ![](https://us1.discourse-cdn.com/flex019/uploads/aerospike/original/1X/395f78cd1e474448adfc2556f25df4d0ab8ded31.png)

There is quite a few events that had been dropped during the writes as I can see from the client\_write\_error.

I am currently having my client do threaded writes and I think that is adding load to the cluster. I am curious to know if there is any parameter than we can set in the aerospike server that can perform threaded writes to the disk.

Thanks, Soudipta

---

<div class="post-metadata">

**Author:** ![kporter](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/kporter/32/515_2.png) [@kporter](https://discuss.aerospike.com/u/kporter)\
**Post date:** [March 31, 2017, 4:56pm UTC](https://discuss.aerospike.com/t/writes-failing-to-aerospike-cluster/4009/6 "2017-03-31T16:56:13Z")

</div>

The value of [fail\_key\_busy](http://www.aerospike.com/docs/reference/metrics#fail_key_busy) would indicate you are hitting one or more hotkeys. Which happens when you have more than [transaction-pending-limit](http://www.aerospike.com/docs/reference/configuration#transaction-pending-limit) transactions queued to a single key.

You should do your best to design your application to avoid hot keys. You can also try (if your use case allows) using the [replace or replace-only](http://www.aerospike.com/apidocs/java/com/aerospike/client/policy/RecordExistsAction.html#REPLACE) record-exists policies which can significantly improve write performance.

---

<div class="post-metadata">

**Author:** ![kporter](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/kporter/32/515_2.png) [@kporter](https://discuss.aerospike.com/u/kporter)\
**Post date:** [March 31, 2017, 4:59pm UTC](https://discuss.aerospike.com/t/writes-failing-to-aerospike-cluster/4009/7 "2017-03-31T16:59:08Z")

</div>

> [@sdas](#):
>
> So if the messages are actually getting dropped, is there any parameter in the Aerospike client that I can use to guard against it.

In this case the client will receive [error code 14 - key busy](https://discuss.aerospike.com/t/aerospike-error-codes/3835) response from the server.

---

<div class="post-metadata">

**Author:** ![pgupta](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/pgupta/32/2351_2.png) [@pgupta](https://discuss.aerospike.com/u/pgupta)\
**Post date:** [March 31, 2017, 5:05pm UTC](https://discuss.aerospike.com/t/writes-failing-to-aerospike-cluster/4009/8 "2017-03-31T17:05:29Z")

</div>

You could also try increasing the transaction-pending-limit. If that reduces the problem, you are on the right track. Increasing the transaction-pending-limit will adversely affect write latency. So not the best way to fix this problem. You need to mitigate the hot key by modeling your data differently.

```
asadm
Admin>asinfo -v 'set-config:context=service;transaction-pending-limit=40'

```

---

<div class="post-metadata">

**Author:** ![sdas](https://avatars.discourse-cdn.com/v4/letter/s/bc79bd/32.png) [@sdas](https://discuss.aerospike.com/u/sdas)\
**Post date:** [March 31, 2017, 8:09pm UTC](https://discuss.aerospike.com/t/writes-failing-to-aerospike-cluster/4009/9 "2017-03-31T20:09:51Z")

</div>

I think I have reduced the hotkey issue by reducing the number of threads on my client while doing the writes. Maybe this is just interim and like you are suggesting remodelling the schema is the correct way to go. But coming back the issue, I still see dropped writes after reducing the number of threads, only this time I dont see fail\_key\_busy count to be increased. Which I assume seems to have solved the hotkey issue if I am not wrong. But looking at the logs, I found out this error:

```auto
Mar 31 2017 20:03:47 GMT: WARNING (drv_ssd): (drv_ssd.c:4147) {user-profile-dev} write fail: queue too deep: q 1513, max 1512. My write-block-size is 1M and max-write-cache is 1512M.

```

---

<div class="post-metadata">

**Author:** ![kporter](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/kporter/32/515_2.png) [@kporter](https://discuss.aerospike.com/u/kporter)\
**Post date:** [March 31, 2017, 8:36pm UTC](https://discuss.aerospike.com/t/writes-failing-to-aerospike-cluster/4009/10 "2017-03-31T20:36:36Z")

</div>

This warning occurs if you exceed the disks throughput capabilities.

See: [https://discuss.aerospike.com/t/why-do-i-see-a-warning-write-fail-queue-too-deep/3009](https://discuss.aerospike.com/t/why-do-i-see-a-warning-write-fail-queue-too-deep/3009)

---

<div class="post-metadata">

**Author:** ![sdas](https://avatars.discourse-cdn.com/v4/letter/s/bc79bd/32.png) [@sdas](https://discuss.aerospike.com/u/sdas)\
**Post date:** [March 31, 2017, 8:45pm UTC](https://discuss.aerospike.com/t/writes-failing-to-aerospike-cluster/4009/11 "2017-03-31T20:45:00Z")

</div>

Yes, was looking at that. Thanks for sharing it. So a number of things I can do here.

1. Decrease concurrency level further - Not a good idea.
2. Remodel my schema so that the same key is not hit multiple times from multiple clients at the same time. I think this is the right way to go.
3. get better aws instances, which I am not so keen on as we will hit the same problem some day or other. In any case we plan to use i3 instances for production use case.

I will update with whichever I go with and how it affects the writes.

Thanks, Soudipta

---

<div class="post-metadata">

**Author:** ![kporter](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/kporter/32/515_2.png) [@kporter](https://discuss.aerospike.com/u/kporter)\
**Post date:** [March 31, 2017, 9:55pm UTC](https://discuss.aerospike.com/t/writes-failing-to-aerospike-cluster/4009/12 "2017-03-31T21:55:31Z")

</div>

> [@sdas](#):
>
> Remodel my schema so that the same key is not hit multiple times from multiple clients at the same time. I think this is the right way to go.

Agreed, if the root problem isn’t addressed it will likely surface again.

---

<div class="post-metadata">

**Author:** ![sdas](https://avatars.discourse-cdn.com/v4/letter/s/bc79bd/32.png) [@sdas](https://discuss.aerospike.com/u/sdas)\
**Post date:** [March 31, 2017, 9:59pm UTC](https://discuss.aerospike.com/t/writes-failing-to-aerospike-cluster/4009/13 "2017-03-31T21:59:02Z")

</div>

Also wanted to ask if there a error code that I can catch for the queue too deep error. Is it 152.

---

<div class="post-metadata">

**Author:** ![sdas](https://avatars.discourse-cdn.com/v4/letter/s/bc79bd/32.png) [@sdas](https://discuss.aerospike.com/u/sdas)\
**Post date:** [March 31, 2017, 10:05pm UTC](https://discuss.aerospike.com/t/writes-failing-to-aerospike-cluster/4009/14 "2017-03-31T22:05:38Z")

</div>

Its 18. Just looked at the docs.

---

<div class="post-metadata">

**Author:** ![kporter](https://sea1.discourse-cdn.com/flex019/user_avatar/discuss.aerospike.com/kporter/32/515_2.png) [@kporter](https://discuss.aerospike.com/u/kporter)\
**Post date:** [March 31, 2017, 10:05pm UTC](https://discuss.aerospike.com/t/writes-failing-to-aerospike-cluster/4009/15 "2017-03-31T22:05:40Z")

</div>

See shared KB article:

> [@Why do I see a warning - "write fail - queue too deep" on server and "Error code 18 Device Overload" on the client](https://discuss.aerospike.com/t/why-do-i-see-a-warning-write-fail-queue-too-deep-on-server-and-error-code-18-device-overload-on-the-client/3009/1):
>
> The client may also report errors in the following form:
> 
> com.aerospike.client.AerospikeException: Error Code 18: Device overload.
