WireGuard at 40 Mpps with DPDK, in Go


• 14 min read
WireGuard at 40 Mpps with DPDK, in Go

Over the last few months, I’ve been spending a lot of time experimenting with high-performance packet processing in Go. A lot of that work turned into packetio, one Go API for sending and receiving packets using DPDK, AF_XDP and Mellanox Direct Verbs.

I’ve built a few things on top of it already. wireblast is a packet generator, and last weekend I wrote about Watch the Net, a fast Internet scanner.

This weekend I finally got around to writing up something I’ve been working on for a little while: WireGuard and DPDK. WireGuard is a fun target because a lot of people run it, there’s an excellent Go implementation in wireguard-go, and packetio is Go too. So the question was pretty simple: what happens if I put packetio underneath wireguard-go and try to run WireGuard over DPDK?

Short answer: it works, and it goes pretty fast.
The more interesting answer is that along the way I ran into a few bottlenecks, and some of the fixes turned out to be useful even outside the DPDK case.

The competition

There are two common WireGuard implementations you’ll run into on Linux today.

Kernel WireGuard is the one most people use. Once a packet enters the WireGuard interface, it stays in the kernel. The kernel routes it, encrypts it, wraps it in UDP, and sends it out the NIC.

wireguard-go implements the same protocol in userspace. Here wg0 is a TUN interface, so plaintext packets cross from the kernel into userspace, get encrypted, and then go back into the kernel through a UDP socket.

For these tests I’m deliberately looking at a difficult case: one WireGuard peer pair. One tunnel between two gateways, with lots of traffic flowing through it.

That matters because one WireGuard peer normally looks like a single UDP flow: same source and destination addresses, same UDP ports. A NIC doing RSS sees that as one flow and tends to put all of it on one receive queue. So even if the machine has 24 cores, one tunnel doesn’t automatically spread itself across them.

The test rig is three identical boxes: AMD EPYC 9275F CPUs with 24 cores, ConnectX-6 Dx NICs at 100 GbE, Ubuntu 24.04, all connected through a switch on one tagged VLAN.

                 WireGuard
box 3 ──► box 1 ═══════════► box 2
  ▲         encrypt            decrypt
  └────────── plaintext ◄────────┘

Box 3 sends plaintext traffic toward box 1. Box 1 encrypts it and sends it through the WireGuard tunnel to box 2. Box 2 decrypts it and sends the plaintext back to box 3, where I count what actually arrived.

First, how fast can the wire actually go?

Since this is a 100 GbE test rig, the max number is 148.8 million packets per second. That’s for minimum-size Ethernet frames, though, and WireGuard makes packets quite a bit bigger.

My smallest plaintext frame is 68 bytes because I’m using a tagged interface. After WireGuard padding, the 16-byte transport header, 16-byte authentication tag, plus UDP, IP and Ethernet, that becomes a 130-byte frame on the wire.

So for the smallest packets in this test, the real ceiling isn’t 148.8 Mpps. It’s about 83.3 Mpps.

plaintext encrypted on wire 100 GbE ceiling
68 B 130 B 83.3 Mpps
512 B 578 B 20.9 Mpps
1442 B 1506 B 8.19 Mpps

So that’s the target. For tiny packets, “line rate” means about 83 million packets per second.

Starting from the bottom

Before adding WireGuard, I wanted a basic reference point for the rig itself.
With plain Linux forwarding between the three boxes, no tunnel and no encryption, I get about 15 million packets per second clean.
That’s useful to keep in mind, because if I keep Linux in the packet path, that 15 Mpps number is already a pretty important ceiling to beat.

Then I turned on WireGuard:


packets/sec
kernel WireGuard 550 kpps
wireguard-go 350 kpps

Those are the first numbers to beat.

One thing jumped out immediately: packet size barely matters. I tried 68-byte packets, 512-byte packets and full-sized packets, and the packet rate stays roughly the same.

That’s interesting because the large packets contain around thirty times more data to encrypt, but the packets-per-second number hardly moves at all.

So whatever is limiting us, it probably isn’t the cipher itself. It looks much more like some fixed cost we pay once per packet.

You can see the effect pretty clearly in bandwidth terms. With full-sized packets, kernel WireGuard is only doing around 6 Gbit/s of plaintext on a 100 GbE NIC. With tiny packets, it’s around 0.25 Gbit/s.

And compared with plain kernel forwarding, kernel WireGuard is about 27× lower on this small-packet workload: roughly 550 kpps versus 15 Mpps.

That’s a pretty big gap. So the next question is obvious: where is all that time going?

One very busy CPU

I was curious where all that time was going. My first assumption was that I’d screwed up the benchmark, which wouldn’t be the first time :)

Then I looked at the CPUs. On the encrypting gateway, one core was at 99%, while almost everything else still had plenty of room. perf pointed straight at:

wg_packet_tx_worker

WireGuard already parallelizes the expensive part. Encryption gets farmed out across CPUs, but packets for a given peer eventually come back through a serialized per-peer transmit path before they go on the wire.

That’s a perfectly reasonable design when you have lots of peers, because different peers can make progress independently. In my test, though, I have one very busy peer.

So the 24-core box is effectively waiting on one serialized part of the pipeline. wireguard-go has a different single-core bottleneck. We’ll get to that in a minute.

Giving wireguard-go the NIC

Now for the interesting part: making packetio work with wireguard-go.
wireguard-go has two really nice interfaces: a conn.Bind for encrypted WireGuard traffic and a tun.Device for plaintext traffic. That’s basically the API boundary.

I implemented both on top of packetio.  This gives wireguard-go packetio-backed versions of those interfaces, plus the boring but necessary Ethernet and ARP bits that come with owning a NIC.

Importantly, wireguard-go still does all the actual WireGuard work: handshakes, key rotation, peer configuration, cryptokey routing and wg(8) configuration. I haven’t changed any of that code.

With completely unmodified wireguard-go, I get about 1.5 million packets per second.

That’s immediately around 4× faster than stock wireguard-go and 2.7× faster than kernel WireGuard, without changing a single line of WireGuard itself. That’s pretty neat!

Why?

The obvious answer is kernel bypass, and yes, that’s part of it. packetio gets rid of a lot of normal kernel work. The NIC DMAs packets straight into memory the application already owns, so there’s no normal socket receive path, no TUN copy, and no trip back through a UDP socket.

In other words, we avoid quite a bit of work moving packets back and forth between the kernel and userspace. But profiling showed something more interesting.

In stock wireguard-go, the kernel was only around a third of the CPU cost, while the actual encryption was around 6%. So simply removing kernel overhead doesn’t really explain a 4× improvement. The important part was where the bottleneck was.

Every outbound packet in wireguard-go eventually passes through a single goroutine, RoutineReadFromTUN, which reads plaintext from the TUN and feeds it into the encryption workers. Those workers are already parallel, and when I looked at the profiler, most of them were simply waiting for work.

So the bottleneck wasn’t the cryptography. It was RoutineReadFromTUN, the one goroutine feeding the crypto workers. That gives the whole pipeline one serial stage. No matter how many encryption workers are available behind it, they can only go as fast as that one goroutine can feed them.

packetio made that serial stage much cheaper, which is how we got to around 1.5 Mpps.
But the serial stage was still there. So the next step was pretty obvious: remove it.

Four patches

Up to this point, wireguard-go itself was completely unmodified. From here on out, I started changing it.

Patch 1: stay on the same core

The first patch adds a small fast-path API. It exposes the current session ciphers, a safe way to reserve nonce ranges, access to the replay check, and enough bookkeeping for wireguard-go to keep owning handshakes, rekeying and timers.

The important part is what happens to the packet.
Instead of receiving a packet on one goroutine, handing it through shared queues, letting another goroutine encrypt it, and eventually handing it to yet another sender, the packet can stay on the core that received it:

receive
  ↓
encrypt
  ↓
transmit

No central reader. No shared encryption handoff. No per-peer sender in the fast path.

That removes a lot of queueing and synchronization and I assume better cache locality, since the packet stops bouncing between goroutines and CPUs.

With that first patch alone, throughput jumps from about 1.5 Mpps to roughly 10 Mpps.

Pretty good! And then things started to fall apart :)
At higher rates, the receiving gateway began rejecting huge numbers of perfectly legitimate packets through WireGuard's anti-replay logic.

Patch 2 and 3: the replay window

Every WireGuard data packet carries a counter. The receiver remembers which counters it has already seen so somebody can't capture a valid encrypted packet and replay it later.

wireguard-go's replay window is 8,128 counters wide. Normally that's plenty.
But now I have 16 cores processing packets from the same WireGuard session in parallel. One core might still be working on counter 10,000 while another has already processed counter 18,500. In this case counter 10,000 isn't a replay. It's just late.

But once it falls far enough behind the highest counter the receiver has seen, it's outside the replay window and gets rejected as too old. That became the next bottleneck.

Patch 2 changes how the replay check is locked. Instead of taking the replay-window lock once for every packet, the fast path takes it once for a whole batch and checks all the packets together.

The logic is the same. The difference is that 16 cores are no longer fighting over the same lock once per packet. That made a big difference to peak throughput, but it didn't solve the deeper problem: the replay window itself was still too small for the amount of parallelism we'd introduced.

So patch 3 widens the window from 8,128 counters to 131,008. That sounds like a huge change, but the state is just a bitmap. It goes from roughly 1 KB to 16 KB per session. Pretty cheap.

I tried removing patch 3 again later just to make sure it was actually doing useful work. Peak performance dropped by about 17%, and the throughput became unstable again as I added queues. So we're keeping that one.

Patch 4: stop touching the timers

At this point, something else floated to the top of the profiler: Timers.

Every batch reports activity back to wireguard-go so keepalives, roaming and rekeying continue to work. At tens of millions of packets per second, that meant multiple cores were touching locks and resetting timers millions of times per second.

These are timers that expire in 5, 10 or 120 seconds. Resetting them millions of times a second seemed a little excessive :)

Patch 4 coalesces that bookkeeping so it happens at most once every 50 milliseconds per peer.

At 16 queues, it barely moves the headline number. At 24 queues, where every core is fighting over the same peer's timer bookkeeping, it matters much more: roughly 34 Mpps without it versus just under 40 Mpps with it.

Here's what the progression looks like on the three-box rig with 16 receive queues:


clean peak
unmodified wireguard-go on packetio 1.5 M 1.7 M
+ patch 1: fast path 9.8 M* 10.2 M
+ patch 2: batched replay check —** 24.4 M
+ patch 3: wider replay window 29.7 M 36.0 M
+ patch 4: quieter timers 29.9 M 37.9 M

'Clean' is the highest rate with under 1% loss, median of three runs; 'peak' is the most that ever came out, whatever the loss.

* 2% loss; patch 1 never reached the <1% "clean" threshold.
** No offered rate stayed below 1% loss.

I like this progression because there wasn't one giant rewrite or some magic crypto optimization. Each bottleneck only became visible after the previous one was removed.

The whole patch series adds about 370 lines to wireguard-go and changes five existing lines.

So, how fast is it?

For 68-byte plaintext packets:


delivered vs kernel
wireguard-go 350 kpps 0.6×
kernel WireGuard 550 kpps 1×
wgpio, unmodified wireguard-go 1.5 Mpps 2.7×
wgpio + patches, 8 cores 19.9 Mpps 36×
wgpio + patches, 16 cores 29.9 Mpps 54×

At 24 queues it peaks just under 40 Mpps, although the 16-queue configuration gives me the better clean number on this rig.

That was fun! And because packetio has multiple backends, I tried those too. On my earlier two-box test.  Direct Verbs landed within half a percent of DPDK. Switching to AF_XDP was slightly more expensive because we still have NAPI kernel work, but still got me about 17Mpps. Pretty cool that it’s so easy to change the driver, and really why packetio exists.

Then I started worrying about TCP

At this point I had something really fast, and I was pretty happy with it. But then I started wondering about packet ordering.

To make one WireGuard tunnel use multiple receive queues,  I sent tunnel traffic using 64 different UDP source ports. That changes the outer 5-tuple enough that RSS on the peer's NIC spreads traffic across queues, which is how we get multiple cores working on the same tunnel.

Without doing that, one WireGuard tunnel lands on one receive queue and effectively one core does all the work. On my system that tops out at around 3 Mpps.

With the source port randomization, I can get into the 30 Mpps range. So the trick matters a lot and I wanted to keep it.

The problem was that my first implementation rotated the source port per packet. That gave me a beautiful spread across all the RSS queues, but it also meant packets from the same inner TCP connection could land on different receive queues at the remote gateway.

For example:

TCP packet 1 → WireGuard UDP source port 51821
TCP packet 2 → WireGuard UDP source port 51822
TCP packet 3 → WireGuard UDP source port 51823
...

At the other side, RSS hashes those packets onto different queues. Different queues are handled by different goroutines, and those goroutines don't necessarily finish at the same time. And there was my packet reordering :(

Most of my testing up to that point had been UDP packet-rate testing, so everything looked great. To see what this actually meant for real traffic, I put a Linux TCP sender and receiver behind the two gateways and ran iperf3.

For one (single) inner TCP stream:


Gbit/s retransmits reorder events
kernel WireGuard 7.9 15,019 4
wgpio, rotating port per packet 10.9 28,107 1,641,020
wgpio, one port per inner flow 12.0 211 27

That made the problem pretty obvious. The fix was about fifteen lines.

Instead of rotating the outer source port for every packet, I hash the inner flow and use that hash to pick the WireGuard source port. That means a given inner flow keeps the same outer 5-tuple, so RSS keeps it on the same receive queue at the other gateway.

You still get parallelism across different inner flows. You just stop spraying packets from the same flow across different CPUs.

With eight TCP streams, throughput went from an erratic 16–46 Gbit/s to 66 Gbit/s. With 32 streams, it reached 89 Gbit/s.

The tradeoff is the usual RSS tradeoff: one elephant flow stays on one receive core. I'm okay with that. That's a pretty normal design tradeoff, and the flow-affine version is what I'm using now.

IFF_MULTI_QUEUE and wireguard-go

There was still one result bothering me. Unmodified wireguard-go on packetio topped out at around 1.5 million packets per second.

I started wondering how much of that was really packetio, and how much was something simpler that I had accidentally fixed along the way.

So I put a normal TUN interface back on the plaintext side. packetio still handled the encrypted traffic from the NIC, but plaintext packets went through a regular kernel TUN again, much closer to how wireguard-go normally runs.

With wireguard-go's usual single-queue TUN, I got about 450 kpps. That looked familiar.

Then I opened the TUN using Linux's multi-queue support and added a reader for each queue. 1.5 Mpps.

Interesting. So I took packetio out completely. Upstream wireguard-go. Normal UDP sockets. Normal kernel networking. The only change was opening the TUN with IFF_MULTI_QUEUE and giving it multiple readers.

The result was the same: 1.5 million packets per second, on about five cores.
That's roughly three times kernel WireGuard on this particular single-peer, small-packet workload, with wireguard-go's protocol and packet-processing code otherwise unchanged.

I think that's one of the most useful results from this whole project.
You don't need DPDK if all you want is to get stock wireguard-go from a few hundred thousand packets per second to around 1.5 Mpps, just stop feeding it through a single TUN queue.

After that, both versions hit the same next bottleneck: RoutineReadFromTUN, wireguard-go's central plaintext reader, at around 1.5 to 1.6 million packets per second.

Same path, same wall. So the lesson is pretty simple: multi-queue TUN gets you to about 1.5 Mpps with otherwise stock wireguard-go. If you want to go beyond that, you need to change the architecture.

That's where the fast path and packetio start to matter, and where the numbers jump from around 1.5 Mpps into the tens of millions.

So when does packetio matter?

This experiment gave me a much cleaner answer than I had when I started. Up to about 1.5 million packets per second, you don't really need packetio. A multi-queue TUN gets stock wireguard-go there with normal UDP sockets and the normal Linux networking stack.

Past that point, though, you have to change the architecture.

The first fast-path patch gets us to roughly 10 Mpps. Once the replay-window and timer bottlenecks are fixed as well, that climbs to around 30 Mpps clean, with peaks just under 40 Mpps.

At that point we're more than an order of magnitude beyond stock wireguard-go.
But there is a tradeoff. The fastest version of wgpio keeps the packet path entirely in userspace. It owns the NIC and forwards packets itself, which is exactly why it can go so fast.

It also means Linux no longer sees that transit traffic. Your normal routing table isn't forwarding it. Your firewall isn't filtering it. tcpdump on the host won't see those packets either.

If you want to keep those Linux features, you can put the plaintext side back through a TUN interface and hand the traffic back to the kernel. That still performs well, but you pay for the extra kernel crossings.

That's really the tradeoff with userland networking: the further you move the dataplane out of the kernel, the more performance and control you get, but the more of the kernel's networking machinery you have to replace yourself.

Caveats

A few things to keep in mind. One peer pair. All of these tests use a single WireGuard peer pair. That's deliberately the difficult case. With hundreds of peers, you get much more natural parallelism.

No GSO or TSO. One frame in, one frame out. I left segmentation offload out on purpose because I was mostly interested in packets per second, and that's where the bottlenecks in this post really show up.

IPv4 only. That's all I tested here.

Flow affinity has a tradeoff. Keeping one inner flow on one outer UDP tuple means one elephant flow also stays on one receive core. That's intentional, and it's a pretty normal RSS-style tradeoff.

Wrapping up

I started this experiment because I wanted to know whether packetio could carry a real application instead of just another packet generator. Now we know it can.

A single WireGuard tunnel gets essentially line rate on a 100 GbE link once the packets get moderately large, and with small packets it gets to around 30 million packets per second, with peaks close to 40 Mpps. For comparison, kernel WireGuard on the same workload did about 550 kpps. So we're talking roughly 54× more packets per second clean, and more than 70× at peak.

But the part I found most interesting wasn't really the headline number. It was how little code it took, and how different the bottlenecks were from what I expected going in.

I assumed the crypto would be the hard part. It wasn't. One goroutine was holding back dozens of encryption workers. A replay window that had always been more than large enough suddenly wasn't. Timer bookkeeping started to matter. And my clever RSS trick for spreading one tunnel across many cores turned out to be terrible for TCP ordering.

We worked through each of those one at a time, and I'm really happy with where it ended up.

The other surprise is that one of the biggest improvements doesn't need packetio at all. Give stock wireguard-go a multi-queue TUN and it goes from a few hundred thousand packets per second to around 1.5 Mpps on this workload.

If you want to go much beyond that, though, you have to start changing the datapath. That's where keeping more of the packet processing in userspace, using something like packetio over Direct Verbs, AF_XDP or DPDK, starts to make a real difference.

Anyway, that's it. You made it to the end :) Thanks for reading.
Let's push some packets, encrypted this time.

Cheers,
Andree

GO TOP