The core thesis should be something like:
I’d been giving customers the same Solr performance advice for years. When I finally isolated and benchmarked each recommendation, I discovered that some of it worked for reasons I hadn’t understood, some of it did essentially nothing, and one test uncovered an unrelated production bug. In the end, a few tiny changes made Sunspot roughly 5–6× faster.
That gives the post three layers simultaneously:
- Surface problem: How fast can I make Rails/Sunspot talk to a small Websolr index?
- Investigation: Which client/configuration changes actually affect wall-clock latency, and why?
- Larger lesson: Performance advice becomes folklore unless you test the mechanism behind it.
The important editorial change is to withhold conclusions until the experiment earns them. In the old post, you explain that Keep-Alive, load balancing, and compression are good, then implement them. In the rewrite, each one starts as a hypothesis.
So instead of:
HTTP Keep-Alive reduces connection overhead.
you want:
One of the things I’d been telling customers forever was to use Keep-Alive. Fine. Let’s see how much it buys us.
Then: nothing.
Then: why the hell did nothing happen?
Then tcpdump.
That pattern should drive the whole article.
Keep the measurement model visible
The QTime/wall-time distinction is central and should appear very early.
It gives you a clean way to establish what you're actually optimizing:
- QTime: roughly, how long Solr reports spending executing the request.
- Wall time: what the application/user actually experiences.
If QTime stays flat while wall time collapses, you have empirical evidence that the gains are outside Solr itself.
That’s much stronger than simply asserting “network overhead matters.”
I’d even establish a small recurring convention in the benchmark tables:
Wall ↓, QTime →
That visually reinforces the point throughout the article.
Treat every test as answering a question
Avoid sections named merely after technologies where possible. Instead of:
- Typhoeus
- Gzip
- JSON Updates
use headings that communicate the hypothesis or surprise:
- Is Net::HTTP the Bottleneck?
- Wait, Why Didn’t Keep-Alive Help?
- Does Sending Less Data Matter?
- JSON Is Smaller. Does Anyone Care?
- Surely Replica Routing Helps
- I Found a Bug. That Wasn’t the Answer Either.
- Can We Just Add More Concurrency?
That keeps the post narrative even while it gets very technical.
Give experiments unequal weight
This is important. You have ten tests/results, but the article shouldn't feel like ten equally sized lab exercises.
The major narrative beats are:
baseline → Typhoeus → Keep-Alive surprise → gzip → routing surprise/bug → final configuration
Those deserve space.
The Ruby response format, JSON writes, and Hydra are supporting experiments. They demonstrate rigor and eliminate plausible alternatives, but they shouldn’t compete with the major discoveries.
I’d probably give wt=ruby a particularly short and funny treatment:
Ruby has a native response writer. Maybe skipping JSON parsing is faster.
It was 2.7× slower.
It also uses
Kernel.evalon a response body from another computer.Moving on.
That is both useful and extremely on-brand.
Make the routing bug a real subplot
Don't bury this as “while testing I found a bug.”
It's the point where the article stops being a straightforward benchmark and becomes an actual debugging story.
The sequence matters:
- You expect replica routing to improve throughput.
- It doesn't.
- That result contradicts your mental model.
- You inspect the routing implementation.
- You find the shared-PRNG/concurrency problem.
- You fix it.
- You confidently retest.
- The path you fixed still doesn't meaningfully improve.
- An adjacent path gets ~21% faster for incidental reasons.
That’s great because the temptation is to declare victory at step 6.
Instead, reality says: yes, you found a bug; no, that doesn't mean the bug caused what you were measuring.
That's probably the strongest epistemological moment in the whole article.
Do not oversell JSON update format
This is actually one of the more useful negative results.
You find:
- JSON is ~42% smaller uncompressed.
- Only ~11% smaller compressed.
- Actual reindex throughput doesn't materially change.
That is a perfect demonstration of optimizing the wrong variable.
The fact that something is objectively “better” in isolation doesn't imply that it matters to system throughput.
Keep it.
Hydra should end in epistemic restraint
Hydra could easily become a tangent about event loops versus threads. Resist that.
The interesting result isn't “Hydra wasn't faster.”
It's:
I could not create conditions on this plan where Hydra's theoretical advantages should dominate without exceeding the plan's actual concurrency limits.
Therefore:
This benchmark cannot tell us whether Hydra scales better in a different environment.
That's the sort of qualification that strengthens the piece rather than weakening it.
Keep implementation details, but move recipe material downward
Readers will still want the configuration that produced the final result. Give it to them.
But preferably after they understand why each line survived.
The final config becomes the payoff:
# approximately:
Typhoeus
gzip
update_format: :json
No explicit Keep-Alive voodoo. No routing header. No giant optimization framework.
The simplicity is the joke.
The ending should revisit your original advice
Don't finish with “here's the final benchmark.”
Finish with an audit of what Past Rob was telling customers:
| Advice | Verdict |
|---|---|
| Use a better HTTP client | Absolutely |
| Enable Keep-Alive | Right goal, wrong prescription |
| Compress payloads | Yes |
| Use JSON updates | Cleaner/smaller, but not responsible for much here |
| Route reads to replicas | Useful operationally, not demonstrated as a throughput win here |
| More concurrency | Depends entirely on downstream capacity |
That makes the post about changing your own understanding, not merely changing some software.
Suggested outline
Title / opening
Working title:
I Took My Own Solr Performance Advice. Some of It Was Wrong.
Subtitle/deck if you use one:
How three small changes made Rails + Sunspot 5–6× faster—and taught me that I didn’t understand my own recommendations as well as I thought.
1. The Sad New Relic Chart
Open with the recurring support interaction.
Customers show up:
“Why is Solr taking 250ms?”
You explain that Solr itself is probably only responsible for a small portion of that. You give your usual recommendations.
Then the turn:
I'd been giving some version of this answer for years.
The embarrassing part was that I'd never sat down and isolated each recommendation to see exactly how much it mattered.
Then establish the challenge:
How much throughput can I wring out of a basic Websolr Hobby index without changing Solr itself?
That’s the hook.
2. What Does “Slow” Mean?
Introduce the three workloads:
- reindex 50K documents
- 1,000 searches
- 1,000 mixed read/write operations
Then explain the two clocks:
Solr QTime vs application wall time.
The experiment is intentionally looking for:
QTime stays roughly constant while wall time changes.
Because you're not tuning queries or indexes. You're tuning everything surrounding Solr.
Establish hardware/plan/index constraints once here so later results have context.
3. The Vanilla Baseline
Rails + Sunspot exactly as shipped:
- RSolr
- Faraday
- Net::HTTP
wt=jsonupdate_format=xml- no special routing
- no explicit compression tweaks
Run all three benchmarks.
Show wall time and QTime.
Then:
Now we start pulling levers one at a time.
4. Can We Beat JSON?
Test wt=ruby.
Explain why it was plausible: native Ruby objects might avoid JSON decoding overhead.
Then reveal:
- parsing is ~2.7× slower
- implementation involves raw
Kernel.eval - RSolr cannot parse XML or Javabin responses anyway
Conclusion:
JSON wins by default and by merit.
Short section. Maybe 3–5 paragraphs.
5. Is Net::HTTP the Real Problem?
Trace the stack just enough:
Sunspot → RSolr → Faraday → Net::HTTP
Explain why Net::HTTP is a sensible conservative default.
Swap in Typhoeus/libcurl.
Benchmark.
Big reveal #1: this is the dominant performance win.
Show QTime unchanged and wall time collapsing.
Now the article has momentum.
6. Wait—Why Didn't Keep-Alive Help?
This should probably be one of the best sections.
You’ve been telling customers to enable Keep-Alive.
So explicitly set the header.
Benchmark.
Nothing.
That's not what I expected.
Investigate with tcpdump.
Discover:
- Net::HTTP: effectively N TCP connections for N operations in your setup
- Typhoeus: reuses the same connection by default
- explicit
Keep-Aliveheader wasn't causing the improvement at all
Reframe the advice:
Connection reuse matters. The header wasn't the mechanism.
That distinction is important.
7. Less Wire, More Speed?
Enable gzip.
Benchmark.
Show the result.
Emphasize where it helps most: larger transfers / tail latency rather than some magical uniform multiplier.
This validates one piece of the original advice.
8. XML vs JSON: Smaller Isn't Necessarily Faster
Switch write/update format.
Measure payload size:
- JSON ~42% smaller raw
- ~11% smaller once compressed
Then benchmark reindexing.
Essentially no throughput difference.
Explain why that result makes sense after observing the system:
server-side indexing and other costs dominate enough that shaving bytes here doesn't move total runtime meaningfully.
This is your cleanest local optimization ≠ system optimization example.
Still keep JSON in the final stack if it is objectively cleaner/smaller and has no downside.
9. Surely Load Balancing Helps
Introduce prefer-replica.
State the tradeoff before measuring it:
more read distribution, but less freshness/NRT consistency.
Measure the actual staleness (~31 seconds).
Then benchmark throughput.
No meaningful improvement.
Again:
That's weird.
10. Oh Look, a Bug
Investigate routing.
Find the concurrency/shared-PRNG bug.
Explain only enough implementation detail to understand why it could serialize or degrade routing behavior.
Fix it.
Deploy.
Now expectations rise.
11. The Bug Was Real. It Wasn't the Whole Story.
Retest.
And get the weird result:
prefer-master: ~21% faster despite being mechanically equivalent to the ordinary/default routeprefer-replica: the path most directly affected by your fix, no meaningful throughput gain
This deserves some reflection.
You absolutely found and fixed a bug.
But your hypothesis that this bug explained the benchmark result was wrong.
Don't force an explanation you don't have.
Something like:
At this point I had reached the least satisfying but most useful answer in performance work: there was clearly more going on than the model I'd started with.
Excellent place to show uncertainty.
12. Fine. What If We Throw Concurrency at It?
Introduce Hydra.
Brief explanation:
- Typhoeus requests traditionally invoked through threads
- Hydra/libcurl multiplexing can manage concurrent transfers efficiently from a single event loop
Test conservatively.
Comparable/underwhelming result.
Then explain the experiment's limitation:
the Hobby plan only permits roughly two useful concurrent connections, so you hit the service ceiling before Hydra gets a chance to demonstrate much.
Conclusion:
No evidence Hydra helps here; no evidence it wouldn't help at larger scale.
Move on.
13. What Actually Survived?
Now assemble the final stack based only on things the experiments justified:
- Typhoeus
- gzip
- JSON updates
Explicitly absent:
- Keep-Alive header
- replica-routing performance hack
- clever concurrency machinery
Then show the table:
| Scenario | Defaults | Best stack | Speedup |
|---|---|---|---|
| Reindex 50K docs | 637.3s | 120.8s | 5.3× |
| 1K searches | 245.6s | 44.8s | 5.5× |
| 1K mixed ops | 335.1s | 55.4s | 6.0× |
Then remind readers:
Solr's QTime barely moved.
That's the punchline to the performance investigation.
Solr didn't become 6× faster.
The system became 6× faster at talking to Solr.
14. I Should Probably Update My Support Macro
Now revisit the original three recommendations.
This is where the post becomes larger than the benchmark.
You began trying to validate:
HTTP Keep-Alive + compression + load balancing.
You ended with:
use a client that actually reuses connections + compress + don't optimize things that aren't bottlenecks.
And somewhere along the way you fixed an unrelated production bug.
The lesson shouldn't be “Typhoeus good.”
It should be:
Performance advice decays into ritual when we stop asking what mechanism it's supposed to affect.
“Use Keep-Alive” sounds precise but actually conflated an objective—avoid repeated TCP/TLS setup—with one supposed implementation.
“Use JSON” sounded faster because the payload was smaller, but size wasn't limiting throughput.
“Use replicas” sounded faster because more hardware should mean more throughput, but that wasn't what your measured workload was bottlenecked on.
Then the ending:
I started this because I wanted some benchmark numbers to back up advice I'd already been giving customers.
I got the benchmark numbers.
I also learned I needed different advice.
That feels like the right final note.
The post starts with expert confidence and ends with better-calibrated expertise. That's the transformation that gets this from an old-school technical blog post to something that fits the newer corpus.