# Multithreaded Crystal initial thoughts

**URL:** <https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089>\
**Category:** Offtopic\
**Created:** [September 3, 2019, 2:17pm UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089 "2019-09-03T14:17:43Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![jgaskins](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/jgaskins/32/2449_2.png) [@jgaskins](https://forum.crystal-lang.org/u/jgaskins)\
**Post date:** [September 3, 2019, 2:17pm UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/1 "2019-09-03T14:17:43Z")

</div>

The multithreading preview for Crystal was [just merged in yesterday](https://github.com/crystal-lang/crystal/pull/8112) so I’ve been playing around with it a bit to check its performance. Sure enough, I was able to saturate all 32 cores on a DigitalOcean droplet with `32.times { spawn { loop {} } }`:

 ![Crystal 32 cores](https://canada1.discourse-cdn.com/flex036/uploads/crystal_lang/original/1X/baec43d0b09fd5c6d630af0e659230ae53a95c9d.png)

I wrote up [a quick `HTTP::Server` app](https://gist.github.com/jgaskins/ba74aaca0ce45f714caaa7571703beba) to check performance (I don’t know how to get a local build of Crystal to use shards or I’d check on a real app). Here’s the single-thread benchmark (via `wrk`, output trimmed):

```auto
  Thread Stats Avg Stdev Max +/- Stdev
    Latency 199.01us 90.71us 8.53ms 97.75%
Requests/sec: 50590.09

```

For comparison purposes, the release version of v0.30.1 gets 50013 reqs/sec with the same code on my machine. Still in the same ballpark, but it’s really heartening to know that the changes didn’t result in worse performance in single-thread mode. In multi-thread mode with a single thread, performance did drop a bit to 47457 — approximately a 5% reduction.

With the `preview_mt` flag enabled (`crystal run -Dpreview_mt --release check_mt.cr`) on Crystal `master`:

```auto
  Thread Stats Avg Stdev Max +/- Stdev
    Latency 93.54us 77.55us 3.36ms 96.54%
Requests/sec: 108131.26

```

This is awesome, we get more throughput from a single process! Latency is lower! In fact, this is the first time I’ve been able to get `wrk` to consume more than 100% CPU before!

 ![Makes the wrk process use > 100% CPU! 😂](https://user-images.githubusercontent.com/108205/64142549-6efab700-cdda-11e9-928d-b6c769c23eb6.png)

It’s not proportional, though, unfortunately. The first version of the app consumes 100% CPU and the second consumes 460%. But rather than a 4.6x improvement in throughput we only get 2.14x. Still a win, just not the one I was expecting. This isn’t intended as a criticism, just an observation that it’s probably not ready yet. :-)

If someone can drop some tips on how to get a locally built Crystal to compile with shards I’d be happy to run this against a real app instead of poorly simulated work. 😂 I have a feeling that throughput might be a little more proportional when the request does real work since scheduling fibers will likely be a much smaller ratio of the total work being done. Getting the DB to keep up might be challenging, though.

---

<div class="post-metadata">

**Author:** ![stakach](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/stakach/32/58_2.png) [@stakach](https://forum.crystal-lang.org/u/stakach)\
**Post date:** [September 3, 2019, 2:38pm UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/2 "2019-09-03T14:38:21Z")

</div>

I feel like the call to spawn should default to the same thread (maintaining backwards compatibility)  
Whereas if you want a new thread you could do something like

```
spawn new_thread: true { operation }
```

---

<div class="post-metadata">

**Author:** ![Blacksmoke16](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/blacksmoke16/32/1241_2.png) [@Blacksmoke16](https://forum.crystal-lang.org/u/Blacksmoke16)\
**Post date:** [September 3, 2019, 2:40pm UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/3 "2019-09-03T14:40:47Z")

</div>

Thats already a thing

[https://crystal-lang.org/api/master/toplevel.html#spawn(\*,name:String?=nil,same\_thread=false,&block)-class-method](https://crystal-lang.org/api/master/toplevel.html#spawn(*,name:String?=nil,same_thread=false,&block)-class-method)

however `same_thread` is defaulted to false.

---

<div class="post-metadata">

**Author:** ![stakach](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/stakach/32/58_2.png) [@stakach](https://forum.crystal-lang.org/u/stakach)\
**Post date:** [September 3, 2019, 8:53pm UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/4 "2019-09-03T20:53:57Z")

</div>

Yeah I know, however I figure that spawning new fibers will still be more common than threads and all existing code expects to be spawning a fiber

---

<div class="post-metadata">

**Author:** ![jgaskins](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/jgaskins/32/2449_2.png) [@jgaskins](https://forum.crystal-lang.org/u/jgaskins)\
**Post date:** [September 3, 2019, 9:07pm UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/5 "2019-09-03T21:07:17Z")

</div>

`spawn` doesn’t create a new thread with this implementation. 😄 In fact, new threads are never created after it begins executing your code. The scheduler spins up a thread pool during bootstrapping and new fibers are assigned to one of those threads. 🤯

It looks like it [assigns them in a round-robin](https://github.com/crystal-lang/crystal/blob/28114b974d49e7638c929ea18d28ed4643e404e6/src/crystal/scheduler.cr#L169-L176) so it may or may not be assigned to the same thread as the fiber that spawned it.

---

<div class="post-metadata">

**Author:** ![asterite](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/asterite/32/60_2.png) [@asterite](https://forum.crystal-lang.org/u/asterite)\
**Post date:** [September 3, 2019, 9:56pm UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/6 "2019-09-03T21:56:30Z")

</div>

> [@jgaskins](#):
>
> It’s not proportional, though, unfortunately. The first version of the app consumes 100% CPU and the second consumes 460%. But rather than a 4.6x improvement in throughput we only get 2.14x. Still a win, just not the one I was expecting.

I think this is expected. Part of the time goes in switching contexts between threads and fibers. I believe there might be more optimizations to improve this situation but I don’t think we’ll get to 4.6x improvement in throughput.

It would be interesting to do a similar benchmark using Go. I know Go does context switches much faster because they need to preserve less amount of registers than we do. And of course they are Google too, so… 😁

---

<div class="post-metadata">

**Author:** ![jgaskins](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/jgaskins/32/2449_2.png) [@jgaskins](https://forum.crystal-lang.org/u/jgaskins)\
**Post date:** [September 4, 2019, 12:00am UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/7 "2019-09-04T00:00:01Z")

</div>

> [@asterite](#):
>
> there might be more optimizations to improve this situation but I don’t think we’ll get to 4.6x improvement in throughput.

Oh, true, sorry. I wasn’t actually expecting 1:1 scaling with CPU time vs throughput. But with 4.6x CPU consumption I think 4x throughput is reasonable, leaving ~10% for additional logistical overhead.

Either way, the simplicity and expressiveness of this implementation is unbelievably good. I believe that’s more important as a starting point. I’m reasonably sure there are some places to optimize that will yield pretty nice — I’ve got my eye on a couple places and I’ll be experimenting with it a bit this week. 🙂

> [@asterite](#):
>
> It would be interesting to do a similar benchmark using Go.

Agreed! I think performance comparisons with Go are awesome to see. Sometimes Go wins, sometimes Crystal wins, and I think that in places where Go is faster there are lessons for Crystal.

---

<div class="post-metadata">

**Author:** ![jgaskins](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/jgaskins/32/2449_2.png) [@jgaskins](https://forum.crystal-lang.org/u/jgaskins)\
**Post date:** [September 4, 2019, 1:47am UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/8 "2019-09-04T01:47:00Z")

</div>

> [@asterite](#):
>
> It would be interesting to do a similar benchmark using Go. I know Go does context switches much faster because they need to preserve less amount of registers than we do. And of course they are Google too, so… 😁

I decided to try it out because I was curious myself. I am not a Go programmer and, in fact, this is literally the first Go program I’ve ever written, but I was able to make it work. I updated the gist I linked in my original post to include that Go app.

 ![Go server using 465% CPU](https://canada1.discourse-cdn.com/flex036/uploads/crystal_lang/original/1X/7d6b5b36601d45eb227ce6058779b5803ee7e80a.png)

The Go server uses the same amount of CPU. Hope you’re sitting down for this next part:

```auto
  Thread Stats Avg Stdev Max +/- Stdev
    Latency 129.40us 91.54us 7.59ms 95.51%
Requests/sec: 69332.77

```

… and Crystal beats it in throughput by 56%. 🤯 Feel free to check my work because, like I said, this is the first Go program I’ve ever written.

---

<div class="post-metadata">

**Author:** ![asterite](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/asterite/32/60_2.png) [@asterite](https://forum.crystal-lang.org/u/asterite)\
**Post date:** [September 4, 2019, 12:51pm UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/9 "2019-09-04T12:51:27Z")

</div>

> [@jgaskins](#):
>
> and Crystal beats it in throughput by 56%

Cool!

I see that in your benchmark you serialize JSON. Go is known to have slow JSON serialization because it uses reflection (and Crystal doesn’t). Could you try running a benchmark where Crystal and Go just send “Hello world” in the response? That way we would be comparing just the HTTP serving part which is where multithreading (context switches, scheduler, etc.) are mainly exercised.

But of course, real world apps use JSON serialization so even without the simpler benchmark this is great news! Thank you for doing these benchmarks ❤

---

<div class="post-metadata">

**Author:** ![jgaskins](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/jgaskins/32/2449_2.png) [@jgaskins](https://forum.crystal-lang.org/u/jgaskins)\
**Post date:** [September 4, 2019, 10:07pm UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/10 "2019-09-04T22:07:38Z")

</div>

> [@asterite](#):
>
> I see that in your benchmark you serialize JSON. Go is known to have slow JSON serialization because it uses reflection (and Crystal doesn’t). Could you try running a benchmark where Crystal and Go just send “Hello world” in the response?

Removing the JSON serialization from the Go app only added ~18% throughput.

```auto
  Thread Stats Avg Stdev Max +/- Stdev
    Latency 109.05us 119.22us 9.19ms 99.46%
Requests/sec: 81592.82

```

> **Golang code here**
>
> ```go
> package main
> 
> import (
> "fmt"
> "log"
> "net/http"
> )
> 
> func main() {
> http.HandleFunc("/", func(w http.ResponseWriter, r *http.Request) {
> fmt.Fprintf(w, "Hello")
> })
> 
> log.Fatal(http.ListenAndServe(":54321", nil))
> }
> 
> ```

It appears Crystal returns a nontrivial JSON payload faster than Go writes a hard-coded string, but let’s check Crystal anyway:

```auto
  Thread Stats Avg Stdev Max +/- Stdev
    Latency 54.06us 23.38us 723.00us 96.85%
Requests/sec: 173259.23

```

I may be going out on a limb here but … ummm … I think we’re good?

The only thing I changed in the Crystal code was to make the `call` method just run `context.response << "Hello"` to match the Go app (also removed the `Fiber.yield` for the same reason).

---

<div class="post-metadata">

**Author:** ![straight-shoota](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/straight-shoota/32/36_2.png) [@straight-shoota](https://forum.crystal-lang.org/u/straight-shoota)\
**Post date:** [September 5, 2019, 10:17am UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/11 "2019-09-05T10:17:44Z")

</div>

Since your HTTP handlers don’t do anything, each request is executed blazingly fast. I suppose the network stack could already be congested. You could try with more and less threads to see if there is a better utilization ratio.

---

<div class="post-metadata">

**Author:** ![asterite](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/asterite/32/60_2.png) [@asterite](https://forum.crystal-lang.org/u/asterite)\
**Post date:** [September 5, 2019, 10:53am UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/12 "2019-09-05T10:53:57Z")

</div>

> [@straight-shoota](#):
>
> You could try with more and less threads to see if there is a better utilization ratio.

To do that, run the program with an env var like `CRYSTAL_WORKERS=8`

---

<div class="post-metadata">

**Author:** ![jgaskins](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/jgaskins/32/2449_2.png) [@jgaskins](https://forum.crystal-lang.org/u/jgaskins)\
**Post date:** [September 5, 2019, 1:55pm UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/13 "2019-09-05T13:55:30Z")

</div>

> [@straight-shoota](#):
>
> You could try with more and less threads to see if there is a better utilization ratio.

Tuning the thread pool was the first thing I tried. :-)

I was using `CRYSTAL_WORKERS=6` for these benchmarks, which is what allowed it to use up to 450-475% CPU. With the raw-string responses, I was only using ~380% CPU and still handling 170k reqs/sec — the Go code used 450% to do the same work. `kernel_task` was hitting 100% CPU on its own handling the I/O, so I don’t think there’s a way I can get \> 170k while running `wrk` (which was at 180% CPU) and the Crystal server on the same macOS box.

I tried on another beefy DigitalOcean droplet because that’s a closer environment to what people will actually be running web servers on. I couldn’t get it to compile Crystal for some reason and I don’t have time to look into it before work, so I’ll have to try again over the weekend.

---

<div class="post-metadata">

**Author:** ![vlazar](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/vlazar/32/31_2.png) [@vlazar](https://forum.crystal-lang.org/u/vlazar)\
**Post date:** [September 6, 2019, 6:55am UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/14 "2019-09-06T06:55:55Z")

</div>

It would be interesting to compare to Go with `fasthttp` instead of standard `net/http`.

Go with `fasthttp` can be much faster:

> **[TechEmpower Web Framework Performance Comparison](https://www.techempower.com/benchmarks/#section=data-r18&hw=ph&test=plaintext&l=zdjvnj-f)**
>
> Performance comparison of a wide spectrum of web application frameworks and platforms using community-contributed test implementations.

The main reason seem to be less memory allocations:

> **[valyala/fasthttp](https://github.com/valyala/fasthttp)**
>
> Fast HTTP package for Go. Tuned for high performance. Zero memory allocations in hot paths. Up to 10x faster than net/http - valyala/fasthttp

But boy I prefer Crystal

> <https://github.com/TechEmpower/FrameworkBenchmarks/blob/master/frameworks/Go/go-std/src/handlers/handlers.go#L197-L205>

---

<div class="post-metadata">

**Author:** ![jgaskins](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/jgaskins/32/2449_2.png) [@jgaskins](https://forum.crystal-lang.org/u/jgaskins)\
**Post date:** [September 6, 2019, 1:33pm UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/15 "2019-09-06T13:33:42Z")

</div>

> [@vlazar](#):
>
> It would be interesting to compare to Go with `fasthttp` instead of standard `net/http` .

Feel free to try it out if you’re interested! You clearly know more about Go than I do. 😄

My own preference is to look at apps that are performing realistic workloads, which is why I serialized JSON to begin with. I’d like to see it do more realistic amount of work within the request, tbh (I’m not interested in how fast I can make an app do nothing 😂), like talking to a DB, cache, etc. I’m just [having trouble getting that working right now](https://github.com/crystal-lang/crystal/issues/8157) so that’ll have to wait a bit longer.

---

<div class="post-metadata">

**Author:** ![asterite](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/asterite/32/60_2.png) [@asterite](https://forum.crystal-lang.org/u/asterite)\
**Post date:** [September 6, 2019, 2:37pm UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/16 "2019-09-06T14:37:50Z")

</div>

> like talking to a DB

Just note that `crystal-db` isn’t prepared for multithreading right now, but will soon be. So if you want to write a benchmark using that you’ll most likely get crashes or similar.

---

<div class="post-metadata">

**Author:** ![bcardiff](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/bcardiff/32/3_2.png) [@bcardiff](https://forum.crystal-lang.org/u/bcardiff)\
**Post date:** [September 6, 2019, 3:32pm UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/17 "2019-09-06T15:32:25Z")

</div>

> [@Blog: Parallelism in Crystal](https://forum.crystal-lang.org/t/blog-parallelism-in-crystal/1101):
>
> A not so short post can be found at: [https://crystal-lang.org/2019/09/06/parallelism-in-crystal.html](https://crystal-lang.org/2019/09/06/parallelism-in-crystal.html) If you want to have a conversation in the twitter land use

---

<div class="post-metadata">

**Author:** ![rogerdpack](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/rogerdpack/32/117_2.png) [@rogerdpack](https://forum.crystal-lang.org/u/rogerdpack)\
**Post date:** [September 10, 2019, 4:25pm UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/18 "2019-09-10T16:25:30Z")

</div>

With the web requests one is it using 100% cpu’s?

---

<div class="post-metadata">

**Author:** ![jgaskins](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/jgaskins/32/2449_2.png) [@jgaskins](https://forum.crystal-lang.org/u/jgaskins)\
**Post date:** [September 11, 2019, 11:58pm UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/19 "2019-09-11T23:58:55Z")

</div>

yes

---

<div class="post-metadata">

**Author:** ![rogerdpack](https://yyz2.discourse-cdn.com/flex036/user_avatar/forum.crystal-lang.org/rogerdpack/32/117_2.png) [@rogerdpack](https://forum.crystal-lang.org/u/rogerdpack)\
**Post date:** [September 13, 2019, 1:13pm UTC](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089/20 "2019-09-13T13:13:12Z")

</div>

It would be interesting to see the output from a sampling cpu profiler for those 100% cpu runs (i.e. “where is it using all that extra cpu” since I guess there’s theoretically 8 cores but it’s only 4x as fast… :) hmm…basically just out of curiosity…

[Next page](https://forum.crystal-lang.org/t/multithreaded-crystal-initial-thoughts/1089.md?page=2)
