If you've been following along, you know I've spent the last few months making packets go fast in Go. We looked at AF_XDP, DPDK, and most recently Mellanox Direct Verbs, all behind one library, packetio. On real hardware that ends the way you'd hope: 148 million packets per second, line rate on 100G, from a handful of Go workers.
This weekend I wondered how much of that survives on EC2. I already knew the DPDK backend worked on Amazon's ENA network interfaces, but I'd only ever tried it on small instances. Meanwhile, AWS sells network-optimized instances with numbers like 300 and 600 Gb/s on the spec sheet. So my question for today: how well do those machines actually hold up when you point a packet generator at them?
Two things I wanted out of this:
- Take packetio for a ride on some seriously beefy machines.
- And this is the bigger one: AWS publishes bandwidth for every instance type, but publishes nothing about packets per second. If you do anything with networking in software, you know packets per second is the number that matters. So let's go measure it.
The instances
I went up the c8gn ladder, Amazon's current Graviton network-optimized family, from the 100G size all the way to the 600G one:
| instance | vCPU | published bandwidth |
|---|---|---|
| c8gn.8xlarge | 32 | 100 Gbps |
| c8gn.12xlarge | 48 | 150 Gbps |
| c8gn.16xlarge | 64 | 200 Gbps |
| c8gn.24xlarge | 96 | 300 Gbps |
| c8gn.48xlarge | 192 | 600 Gbps |
These are not cheap machines, but the specs are pretty impressive! The 48xlarge pair costs about $23 an hour, so every test is designed to be short and to get the answer in one go.
How I tested
The setup was identical for every instance type:
- A pair of identical instances, same VPC, same subnet, same availability zone, talking to each other over a dedicated ENA NIC. Management traffic goes over a separate interface so nothing but test traffic touches the counters we care about.
- DPDK on both ends. The data interface is bound to
vfio-pciand owned by a DPDK application; the kernel never sees a packet. I went with DPDK rather than the kernel or AF_XDP because I wanted the NIC to be the only variable. - packetio's stock examples. The sender is
blast, the receiver isdrop, the same programs from the packetio repo that I used for the hardware numbers. No special EC2 build. Both are Go programs using the DPDK backend, so what you see below is what a Go application would see. - Three packet sizes. 64 bytes for the packet rate, 1500 bytes for what a normal MTU gets you, and 9000 bytes to validate the published bandwidth.
- A queue ladder. For every instance I started with one queue on one core and doubled up until the number stopped moving. That's how you tell a per-core limit from an instance limit.
- Read the results at the receiver. The sender's transmit counter is not what gets delivered, as AWS's shaper sits after it. On one instance the sender happily reported 3.6 Mpps while 1.0 Mpps arrived. Every number here is what the other side actually received.
Two things kept me honest. Every transmit test was repeated with a ~120 line plain C DPDK sender that shares nothing with packetio, and the two always landed within a few percent of each other. And I kept a close eye on the ENA driver's allowance counters, pps_allowance_exceeded, bw_in_allowance_exceeded and friends, which the Nitro card exposes to tell you which limit you just hit. This gives me confidence the numbers aren't implementation specific. I added support for reading those from a DPDK-owned interface to packetio along the way, because without them you're guessing.
Packets per second is everything
The thing about networking in software: every packet costs roughly the same amount of work regardless of how big it is: a lookup, a header rewrite, a descriptor. A router, a firewall, a VPN gateway, a load balancer, all of them are packet per second machines first and bandwidth machines second. Bandwidth is just packets times size. So when a spec sheet says 100 Gbps, the question I actually need answered is: at what packet size?
At 64 bytes, 100 Gbps is 148.8 million packets per second. That's what the ConnectX-6 in my lab does, and it's our benchmark.
One queue, one core
Let's start small. One transmit queue, driven by one core, 64 byte packets:
| instance | one queue, packetio | one queue, plain DPDK |
|---|---|---|
| c8gn.8xlarge | 2.19 Mpps | 2.33 Mpps |
| c8gn.12xlarge | 1.52 Mpps | 1.61 Mpps |
| c8gn.16xlarge | 1.52 Mpps | 1.61 Mpps |
| c8gn.24xlarge | 1.52 Mpps | 1.61 Mpps |
| c8gn.48xlarge | 1.47 Mpps | 1.60 Mpps |
About 1.5 to 2.2 million packets per second per core. On the ConnectX-6 in my lab one core does around 60 million. That's a 30x gap, and it's not packetio: the plain C sender lands in the same place. Bigger batches, deeper rings, prebuilt frames, checksum offload: I tried them all and none of it moves the number.
I don't have a definitive answer for why, but I have a suspect. ENA runs its transmit queues in what the driver docs call Low Latency Queue mode: instead of the NIC fetching your packet from memory, the driver pushes the descriptor and the first 96 bytes of every packet straight into the device over PCIe. For a 64-byte packet that's the whole thing, written by the CPU, one packet at a time. At 2.8 GHz, 2 Mpps works out to roughly 1,450 cycles per packet, against about 50 on the Mellanox card, and that's the kind of gap a PCIe write per packet would explain. You can't test the theory by turning it off: with enable_llq=0 the queue refuses to start, and DPDK's own docs say disabling it is "highly not recommended" anyway.
So on EC2, a core buys you about two million packets per second. The interesting question is what happens when you add more of them.
Adding queues
Same test, now with N transmit queues on N cores, receiver side:
| queues | 8xlarge | 12xlarge | 16xlarge | 24xlarge | 48xlarge |
|---|---|---|---|---|---|
| 1 | 2.19 | 1.52 | 1.52 | 1.52 | 1.46 |
| 2 | 4.39 | 3.08 | 3.08 | 3.08 | 3.06 |
| 4 | 9.12 | 6.33 | 6.33 | 6.33 | 6.29 |
| 8 | 16.51 | 13.34 | 13.35 | 13.34 | 13.25 |
| 16 | 21.12 | 27.71 | 27.75 | 27.73 | 27.51 |
| 24 | 17.15 | 36.65 | 36.65 | 36.65 | 35.22 |
| 32 | 16.47 | 36.66 | 36.67 | 36.67 | 35.26 |
| 48 | - | 36.38 | 36.71 | 36.71 | 35.26 |
| 64 | - | - | 36.48 | 36.73 | 35.34 |
It scales linearly, about 1.6 to 2.2 Mpps per queue, right up until it doesn't. The 8xlarge peaks at 16 queues and 21 Mpps and then actually gets worse with more. And the 12xlarge, 16xlarge and 24xlarge all stop at the same place: 36.6 million packets per second, at 24 queues, flat as a table from there to 64. Adding 40 more cores changes nothing. That's not a CPU limit, that's a wall, and the ENA counters confirm it: push the offer past ~36 M and pps_allowance_exceeded starts ticking on the sender.
The takeaway: on a c8gn.12xlarge or anything bigger, one network interface tops out at roughly 36 million packets per second. It doesn't matter how many queues you use, how many cores you throw at it, or whether the spec sheet says 150, 300 or 600 Gbps. 36M pps, that's the number I was looking for!
The results
Here's the complete table I wish AWS would publish. Everything measured at the receiver, 64 byte packets for the packet rate, 1500 and 9000 byte packets for bandwidth:
| instance | vCPU | published | pps @ 64B | Gbps @ 1500B | Gbps @ 9000B | % of published, at 64B |
|---|---|---|---|---|---|---|
| c8gn.8xlarge | 32 | 100 Gbps | 21.1 M | 102 | 101 | 14% |
| c8gn.12xlarge | 48 | 150 Gbps | 36.6 M | 153 | 151 | 16% |
| c8gn.16xlarge | 64 | 200 Gbps | 36.6 M | 205 | 203 | 12% |
| c8gn.24xlarge | 96 | 300 Gbps | 36.6 M | 282 | 304 | 8% |
| c8gn.48xlarge | 192 | 600 Gbps | 35.3 M | 282 | 304 | 4% |
A few things jump out.
The bandwidth numbers are honest. With 9000 byte packets every instance hit its published number, usually a few percent over. If you move big packets, you get what you paid for.
Packets per second are a different story. 36.6 million packets per second is about a quarter of what a single 100G NIC does in my lab, on an instance sold as 150, 200 or 300 Gbps. At 64 bytes you're only using between 8 and 16 percent of the bandwidth you're paying for. That's not a complaint, it's how the platform works, but it is the number to design around, and it's not on any spec sheet.
There's one packet rate from the 12xlarge up. 150G, 200G, 300G, one network card of the 600G: all ~36 Mpps. The extra bandwidth and the extra cores buy you nothing for small packets. Per gigabit that's 244 kpps on the 12xlarge, 183 on the 16xlarge and 122 on the 24xlarge. If your workload is small packets, the 12xlarge is the sweet spot of this family, and the bigger sizes only make sense when the packets get big too.
Where does bandwidth take over? Since the packet rate is flat, the bandwidth you actually get is simply the packet rate times the packet size, until you hit the published number. I ran a size ladder to find where that happens. The 8xlarge fills its 100G at about 640 bytes. The 12xlarge fills 150G already at 512 bytes. The 16xlarge needs about 700 bytes. The 24xlarge doesn't get there at 1500 bytes at all: 282 Gbps, and that's the sender's NIC topping out, not AWS. Same for the 48xlarge. Above 200G, you need jumbo frames to see the bandwidth you're paying for.
600 Gbps needs two network cards
The 48xlarge was the one I was most curious about, and it came in at exactly the same numbers as the 24xlarge: 35 Mpps, 282 Gbps at 1500 bytes, 304 Gbps with jumbo frames. Half the published bandwidth, and this time with bw_out_allowance_exceeded firing on the sender. AWS was shaping me at 300G on a 600G instance.
The reason: to get 600 Gbps on the 48xlarge you need to use two network cards, each with its own 300 Gbps allowance. Normally the published bandwidth is per instance and every ENI shares it; here it's per card. To get 600 you need two interfaces, one attached to each card (--network-card-index 1 for the second), and something spreading traffic over both. There's no bonding layer doing that for you. Attach one interface, run iperf, and you'll measure a 300G machine.
In hindsight this makes sense. The fastest NICs you can commonly buy today are 400G; anything beyond that is two of them. It would have been odd for AWS to have a single 600G device.
It is documented, sort of. Someone on Twitter kindly pointed me at the network specifications table for compute optimized instances, which has a "Network cards" column, and the c8gn.48xlarge says 2. Not a page you naturally end up on, and nothing there says one interface gets half.
It probably also means the packet rate in the table is per card, so both cards together would be good for ~70 Mpps. I didn't test that (the pair was expensive enough for one weekend), so treat it as a well-founded guess.
What I didn't test
AWS documents a burst and baseline mechanism for bandwidth on the smaller instance sizes. The c8gn sizes I tested all have a fixed bandwidth, and my runs were short by design, a few minutes each, so I have nothing to say about credit buckets here. I did keep the allowance counters in view the whole time and saw no sign of a packet-rate bucket in any run. If you run these instances at full blast for an hour, tell me what you see.
Things fixed along the way
As usual, running on new hardware finds bugs. packetio's DPDK backend now builds on arm64 (all of these are Graviton machines, and the mempool header alignment differs from x86), it can read the ENA allowance counters through DPDK's xstats, and the CPU pinning code learned about hyperthreads: on an Intel instance it used to place one worker per physical core and leave half the vCPUs idle. All of that is in the repo.
Wrap-up
So, how fast can you go on AWS?
- Bandwidth: as fast as it says on the box. Every published number held, as long as the packets are big enough, and above 200G that means jumbo frames. On the 48xlarge it also means two network interfaces.
- Packets per second: about 36 million per network interface, from the 12xlarge up. That's the number AWS doesn't publish, and it's about a quarter of what one 100G NIC does on bare metal.
- Per core: about 2 million pps. Roughly 30x more expensive per packet than a ConnectX. You need 24 cores just to push enough packets to find the instance limit.
So we got the numbers we were after, and packetio got better along the way. If you're building network functions on EC2, you now have the figure that was missing, and it's the one to pick your instance by: know your packet size first. For small packets the budget is one flat number across a 4x range of prices, and a single 100G ConnectX in a bare-metal box does four times the packets of a c8gn.24xlarge for a lot less money.
One more catch, and it has nothing to do with packets. Everything here ran inside one availability zone, where traffic is free. Cross an AZ and it's $0.01 per GB each direction. 300 Gbps is 135 TB an hour, so a 48xlarge at its rated bandwidth costs about $1,350 an hour in transfer, on top of $11 for the instance. Serious networking on AWS needs a serious credit card, and it's the traffic that runs it up. That's a post of its own.
The tooling is in the packetio repo; the blast and drop examples are all you need to reproduce any row in the table. If you run this on an instance family I didn't, I'd love to see the numbers.
That's it, thanks for reading! Pushing packets on AWS is now a little clearer :)
Cheers
-Andree