Paul's Programming Notes PostsRSSGithub

Thrift Is More Difficult To Use Than HTTP

The microservices at my work implement both HTTP endpoints and Apache Thrift RPC endpoints, with Thrift carrying the internal communication between services. External access goes through an API gateway that needs HTTP anyway. I keep losing hours to a Thrift problem I could have solved in minutes over HTTP.

New services don’t get Thrift support at all anymore. They’re documented with Swagger and validated with JSON schema instead, and the tests check requests and responses against the spec.

What makes it harder to live with than HTTP:

  • An exception you didn’t declare in the IDL reaches the client as TApplicationException: Internal error and nothing else. The generated processor catches it, logs the traceback on the server, and sends back that one opaque message, so every debugging session starts with going to find the server log.
  • With an HTTP endpoint I can mock things out with a library like responses. Nothing equivalent exists for Thrift in Python yet, so testing is a lot more work.
  • Version mismatches are very difficult to debug, especially when someone changes the type of an existing field or adds a field to the end of a definition.
  • Updating one endpoint means updating three things: the Thrift definitions, the definitions on the client, and the definitions on the server.
  • The Python tooling for running a Thrift service is nowhere near as mature as it is for HTTP services.
  • A new developer has definitely used HTTP and probably hasn’t used Thrift. The business intelligence people don’t touch it at all, the barrier to entry is too high.
  • Javascript and iOS support isn’t great, though you probably shouldn’t be exposing Thrift services to the public internet anyway.

Thrift does buy real things. It’s strongly typed, the definitions give you one place to look at all of your models, it validates them for you, and the leaner transport puts less over the wire.

That last one matters less than it sounds. Gzipped JSON is already pretty compact and the default Thrift transports don’t compress at all, so you’re saving a few bytes in exchange for everything above.

If you’re only using Python, marshmallow covers the validation, or you can pair JSON schema with something like warlock to build objects from it. If you do stay on Thrift, thriftpy is a big quality of life improvement over the built-in client because it reads the definitions directly instead of making you generate code from them.

Whether the complexity is worth the performance depends on your scale. Uber runs Thrift across a thousand services and Matt Ranney still summed it up as “Thrift is OK, but generated code is bad” in What I Wish I Had Known Before Scaling Uber to 1000 Services, which is the same complaint that makes thriftpy worth using. For a small team it’s a lot of work and learning to end up somewhere HTTP already is.

JSON-API - Lessons Learned

A few things I’ve learned while building against JSON-API:

  • Objects referred to by relationships all go into one shared included array rather than being nested under the relationship that points at them. Without a JSON-API client library that’s a slight pain to parse, because you’re matching type and id pairs back to entries in a flat list. It beats duplicating the same object under every relationship that refers to it, but I would have preferred each object type under its own top level key.
  • It’s a lot more verbose than a response you’d shape by hand. Every record carries type, id, attributes, and relationships wrappers around what would otherwise be a flat object.
  • There’s no PUT. The spec uses PATCH for updates, so you send only the fields that changed instead of replacing the whole resource.

Sphinx Search - Lessons Learned

Here are a few things I’ve learned while working on a project that uses Sphinx search:

  • It’s important to know the difference between fields and attributes. Attributes are basically unindexed columns and you should try to avoid filtering only on these columns. Fields support full text search.
  • It supports its own custom binary protocol and the MySQL protocol (recently they also added a HTTP API). When you see “listen = localhost:9306:mysql41” in the config, that means it’s listening for MySQL protocol traffic on port 9306.
  • https://github.com/a1tus/sphinxapi-py3 appears to be the best Python client for the binary api at the moment. This doesn’t support INSERTing things into the index (you’ll need to use the MySQL protocol for that).
  • The version of sphinxapi-py3 on pypi is a fork with just a few minor fixes and appears to be safe.
  • It does not match partial words by default. Turning on partial matching can also increase the size of your index dramatically. You can also limit the fields that support partial matching with the infix_fields and prefix_fields setting.
  • Stemmers aren’t turned on by default. So, searching for “dog” will not match “dogs”.
  • Most special characters ($, @, &, etc) are ignored by default. You will need to add them to charset_table if you want them to be searchable.
  • Ruby’s thinking-sphinx looks much more battle tested than all of the Python binary api clients: https://github.com/pat/thinking-sphinx
  • You will need to use a real-time index if you want to INSERT/DELETE records immediately.
  • If you’re using a real-time index, you will probably need to increase the rt_mem_limit from its default of 128mb. If this limit is too low, you’ll see a high number of “disk chunks” when you run the “SHOW INDEX rtindex STATUS” query. More info: http://sphinxsearch.com/blog/2014/02/12/rt_performance_basics/
  • You have to use a special dialect if you want to use SQLAlchemy with sphinx: https://github.com/conversant/sqlalchemy-sphinx
  • This appears to be the best Dockerfile for sphinx: https://github.com/leodido/dockerfiles

I probably won’t be using Sphinx search for any new projects. Elasticsearch seems preferable these days.

MySQL - Duplicate Errors & Trailing Whitespace

I had a unique constraint on a VARCHAR column and I inserted two rows with the following values:

  1. “name” (without trailing whitespace)
  2. “name “ (with trailing whitespace)

To my surprise, I got a duplicate error on that 2nd insert. It turns out that MySQL ignores that trailing whitespace when it makes comparisons.

The MySQL docs say this: “All MySQL collations are of type PAD SPACE. This means that all CHAR, VARCHAR, and TEXT values are compared without regard to any trailing spaces. ‘Comparison’ in this context does not include the LIKE pattern-matching operator, for which trailing spaces are significant.” (https://dev.mysql.com/doc/refman/5.7/en/char.html)

The solution? You should probably be trimming trailing whitespace in your API endpoints and on your front-end.

Gevent + Requests Performance With verify=True/False

If you use gevent with requests.get on a HTTPS URL with the default verify=True enabled, you’ll see almost 2x longer execution times than with verify=False.

I made a script to test:

Here are the results:

verify=True took: 40.3454630375 secs verify=False took: 39.3803040981 secs gevent verify=True took: 2.23735189438 secs gevent verify=False took: 1.58263015747 secs

I suspect that gevent is having trouble using pyopenssl concurrently because it’s a C library.

Backing up or dumping a memcached server

I was needing to move from an old cache server to a larger one, but I wanted to do it without flushing cache.

The first thing I came across was this “memcached-tool” which has a dump command: https://github.com/memcached/memcached/blob/master/scripts/memcached-tool

There’s another article that mentions using memdump and memcat: How to dump memcached key/value pairs fast (archived)

Unfortunately, those methods only dumped a few mb of data. This post explains why: https://stackoverflow.com/a/13941700

You can only dump one page per slab class (1MB of data)

So, I ended up writing a script that loops through the expected cache keys, gets the data in cache, then sets the data in the new cache server.

Gunicorn - "Resource temporarily unavailable"

Updated 2026-08-08: explained why gunicorn’s own backlog setting doesn’t fix this.

Are you seeing this error in your logs while your server is under high load?:

[error] 10#0: *14843 connect() to unix:/tmp/gunicorn.sock failed (11: Resource temporarily unavailable) while connecting to upstream, client: 192.0.2.10, server: , request: "GET / HTTP/1.0", upstream: "http://unix:/tmp/gunicorn.sock:/", host: "198.51.100.20"

I ended up making an example dockerfile with nginx + gunicorn + flask to reproduce this problem: https://github.com/pawl/somaxconn_test

Bumping the net.core.somaxconn setting ended up fixing it.

Error 11 is EAGAIN, and on a connect() to a unix socket it means the listening socket’s accept queue is full. Connections sit in that queue after the kernel accepts them and before gunicorn calls accept(), so it fills up whenever requests arrive faster than the workers drain them. Once it’s full the kernel refuses new connections instead of queueing them, and nginx reports the refusal as this error.

net.core.somaxconn is the ceiling on how deep that queue is allowed to be. Linux capped it at 128 until kernel 5.4 raised the default to 4096, so on anything older this is a low bar to hit.

The part that cost me the most time is that gunicorn’s own --backlog defaults to 2048, which looks like plenty. listen(2) silently truncates whatever a process asks for down to somaxconn, so gunicorn requested 2048 and got 128, with nothing in any log to say so. Raising the sysctl is what actually changes the queue:

sudo sysctl -w net.core.somaxconn=4096

Put it in a file under /etc/sysctl.d/ to survive a reboot. In a container it’s a property of the network namespace rather than the image, so it’s docker run --sysctl net.core.somaxconn=4096, or set on the host if the container shares its network namespace.

Worth saying that a full accept queue is usually a symptom. If the workers can’t keep up, the queue depth buys headroom for a traffic spike, not for a slow application.