2026-09-28 20:38:25

The simdjson library is a C++ library to parse and generate JSON. It is used in Node.js, ClickHouse, Meta Velox, StarRocks, Apache Doris, the Ladybird browser and many other systems. We released version 4.0 a year ago, in September 2025. Its headline feature was C++26 static reflection: you could turn a C++ structure into JSON, and back, without writing any glue code.
Today we are releasing version 5.0.
Static reflection is no longer guarded and is an officially supported feature. When your compiler has reflection enabled (e.g., g++ -std=c++26 -freflection with GCC 16), simdjson detects it by itself and the reflection-based functions become available.
When deserializing a C++ structure through reflection, simdjson now uses key selectors (see below) by default: it reads the object in a single pass, whatever the order of the keys. You can return to the previous approach (one lookup per member) with -DSIMDJSON_DISABLE_KEY_SELECTOR_REFLECTION=1.
A positive integer in [2^64, 10^20) is now reported as a big integer (BIGINT_NUMBER), like other overflowing integers, instead of a malformed number. When big integers are parsed as strings, a token such as 123456789123456789123x is now rejected.
One major change is the key selectors. A common task is to extract a few fields from a JSON object. In simdjson 5.0 (C++20 or better), you can name the keys at compile time and visit the object once:
using namespace simdjson;
auto json = R"({ "name": "Daniel", "age": 42, "city": "Montreal" })"_padded;
ondemand::parser parser;
auto doc = parser.iterate(json);
std::string_view name, city;
uint64_t age = 0;
auto result = doc.get_object().for_each<"name", "city", "age">(name, city, age);
// name == "Daniel", city == "Montreal", age == 42
The keys can appear in any order. At compile time, simdjson builds a perfect hash function for your set of keys. At run time, recognizing a key takes a hash computed from a couple of bytes and one comparison. You can also pass one callback per key instead of variables. The iteration stops as soon as all keys have been found.
We added annotations for the data structures for automated (C++26) serialization and deserialization: rename, rename_all, alias, skip, default_value, flatten, deny_unknown_fields, transparent, and so forth.
struct [[= simdjson::rename_all<simdjson::case_style::camel_case>]] User {
std::string first_name;
int64_t user_id;
[[= simdjson::rename<"KEY">]] int api_key;
};
// {"firstName":"Ann","userId":7,"KEY":8}
We often get many JSON documents in one file or one network message. simdjson has long supported streams of documents separated by white space (NDJSON). In 5.0 we added:
stream_format::newline_delimited: you promise that each document sits on its own line. When you only read part of a document, simdjson jumps to the next line instead of walking over the rest of the document.simdjson::slice_at, which cuts a stream into blocks at document boundaries so you can parse the blocks on as many threads as you like. The built-in threaded mode uses at most two threads.We also fixed several bugs in document_stream, found in an audit by Francisco Geiman Thiesen.
There are many other smaller features.
NaN or Infinity, but many systems produce them anyway. If you define SIMDJSON_ENABLE_NAN_INF, simdjson parses them, and serializes them.get_uint8(), get_int8(), get_uint16(), get_int16() check the range for you. With C++23, get_float32() and get_float64() return std::float32_t and std::float64_t. The binary32 value is rounded once, directly from the decimal string, not through a double.parser.parse_unpadded(...)). It is slower than the regular function, but it never reads past the end of your buffer and it does not copy your data.simdjson::padded_input adds padding only when it is needed: when your string ends near a page boundary.std::views::transform and other adaptors.rbegin(), rend()) with no allocation.get_current_position() and revert_position(): if you miss an optional field, you can go back to where you were instead of rescanning the whole object.char8_t (u8) variants of the string accessors in C++20.The simdjson 5.0 release improved performance compared to simdjson 4.0 in some key cases. Let me review some of them.
I built both versions with GCC 16.1 (-O3, CMake Release) and ran them on an Intel Xeon Gold 6548N (Emerald Rapids), pinned to one core.
Let us start with DOM parsing of our standard files (GB/s):
| file | 4.0 | 5.0 | speedup |
|---|---|---|---|
| 4.78 | 4.82 | 1.0 | |
| citm_catalog | 4.82 | 4.79 | 1.0 |
| github_events | 5.33 | 5.30 | 1.0 |
| canada | 1.10 | 1.21 | 1.1 |
| marine_ik | 1.25 | 1.38 | 1.1 |
| mesh | 1.17 | 1.26 | 1.1 |
| numbers | 1.13 | 1.41 | 1.3 |
| twitterescaped | 1.59 | 2.85 | 1.8 |
| update-center | 3.96 | 3.84 | 1.0 |
| apache_builds | 4.95 | 4.78 | 1.0 |
Files full of numbers (canada, marine_ik, mesh, numbers) are 8% to 25% faster. And a file full of escaped Unicode characters (twitterescaped) is almost twice as fast: among other changes, we now decode consecutive uXXXX sequences without going back to the string scanner between them.
We also serialize faster. Printing floating-point numbers used to be a bottleneck. We replaced the ancient Grisu2 by Dragonbox, and removed calls to memcpy and memmove from the hot path.
| file | 4.0 | 5.0 | speedup |
|---|---|---|---|
| 0.94 | 0.96 | 1.0 | |
| citm_catalog | 1.07 | 1.08 | 1.0 |
| gsoc-2018 | 1.10 | 1.25 | 1.1 |
| canada | 0.31 | 0.52 | 1.7 |
| marine_ik | 0.28 | 0.37 | 1.3 |
| mesh | 0.34 | 0.47 | 1.4 |
| numbers | 0.32 | 0.49 | 1.6 |
The simdjson library is a community project. Since version 4.6, contributions came from fior512, 吴杨帆, Alecto Irene Perez, Francisco Geiman Thiesen, Max Bachmann, MoonFlowww, Advit Arora, Jaël Champagne Gareau, Taimoor Kiani, Vasily Pelikh, jmestwa-coder, liyinlong, AlbertoFVisconti, Aylin Dmello, Cuda Chen, Ezra Li, Madhurendra Purbay, Makkar, Pastoray, Paul Dreik, Pavel Kruglov, Piotr Kubaj, Yusuf İhsan Görgel, metsw24-max, neil, pratap singh, Vladimir Saraikin, wankun, xaldarof, Riyane El Qoqui, Justin Li and others. Thank you!
2026-09-27 00:53:53

C# strings are UTF-16. Most characters are one 16-bit code unit. Characters outside the basic multilingual plane, emoji included, take two: a high surrogate (U+D800 to U+DBFF) followed by a low surrogate (U+DC00 to U+DFFF). A surrogate with the wrong neighbor, or with none, is ill-formed.
You should never send an ill-formed string to disk or to the network. It is a bad practice.
In JavaScript, we have fast functions to fix strings or check whether they need fixing:
String.prototype.toWellFormed() replaces every lone surrogate with U+FFFD.isWellFormed() reports whether any replacement is needed.I added both functions to my C# library SimdUnicode, in pull request 54. The algorithm is the same that we contributed to the JavaScript engine V8, so Chrome already fixes strings this way.
string s = UTF16.ToWellFormed(input); // same instance, when the input is already well formed
bool ok = UTF16.IsWellFormed(span);
When the input is well formed, ToWellFormed returns it as is. No allocation.
How are strings fixed? Basically, you replace bad inputs by the replacement character U+FFFD.
Our processors have special instructions called SIMD that allow data parallelism: you can compare multiple values at once. Recent x64 processors from AMD and Intel have better data parallelism than ARM chips, although both have powerful instructions.
The conventional approach in C# to repair a string is a function such as the following.
static void Repair(ReadOnlySpan<char> input, Span<char> output)
{
input.CopyTo(output);
int i = NextError(output, 0);
while (i >= 0)
{
output[i] = 'uFFFD';
i = NextError(output, i + 1);
}
}
// Index of the next lone surrogate at or after 'start', or -1 if none.
static int NextError(ReadOnlySpan<char> s, int start)
{
int i = start;
while (true)
{
int k = s.Slice(i).IndexOfAnyInRange('uD800', 'uDFFF');
if (k < 0) return -1;
i += k;
if (char.IsHighSurrogate(s[i]) && i + 1 < s.Length && char.IsLowSurrogate(s[i + 1]))
i += 2; // valid pair, skip it
else
return i;
}
}
In SimdUnicode, I also use data parallelism.
Let me measure.

On the Xeon, with AVX-512, Latin validates at 69 GB/s against 33 GB/s for IndexOfAnyInRange. The Emoji input is well formed, and it is nothing but surrogate pairs. The runtime search drops to 0.4 GB/s. Our check holds 53 GB/s.

Our results are similar on the M4 Max, although a bit less impressive compared to the Intel results.
Validation can return at the first lone surrogate. The buffer form of ToWellFormed writes every code unit, a copy of the input or U+FFFD. When the input is well formed, it is effectively a memory copy. Thus we can compare the performance against a copy.


Roughly speaking, we are consistently about as fast as a copy.
Versions used: .NET SDK 10.0.400 on Linux, 10.0.103 on macOS. Intel Xeon Gold 6548N (Emerald Rapids). Apple M4 Max.
Clausecker, R., & Lemire, D. (2026). Fixing ill-formed UTF-16 strings with SIMD instructions. Software: Practice and Experience. (arXiv)
2026-09-26 10:48:13

2026-09-25 09:30:49

We create new strings all the time. How quickly can you produce short strings in your programming languge? To create a meaningful string, we convert an integer to a string. In Python, that is str(i). The loop stores each new string in a small ring buffer of 1024 slots, so that the engine cannot simply discard the work.
def from_int(n):
b = buf
for i in range(n):
b[i & 1023] = str(i)
In JavaScript (Node.js and Bun), I use String(i):
function from_int(n) {
for (let i = 0; i < n; i++) buf[i & 1023] = String(i);
}
In C++, I use std::to_string:
void from_int(uint64_t n) {
for (uint64_t i = 0; i < n; i++) buf[i & 1023] = std::to_string(i);
}
In Rust, I use the standard to_string() and, as an alternative, the popular itoa crate, which is designed for fast integer formatting:
fn std_to_string(buf: &mut [String], n: u64) {
for i in 0..n {
buf[(i & 1023) as usize] = i.to_string();
}
}
fn itoa_to_string(buf: &mut [String], n: u64) {
let mut b = itoa::Buffer::new();
for i in 0..n {
buf[(i & 1023) as usize] = b.format(i).to_owned();
}
}
In Go, I use strconv.Itoa:
func fromInt(n int) {
for i := 0; i < n; i++ {
buf[i&1023] = strconv.Itoa(i)
}
}
In Nim, I use the $ operator:
proc fromInt(n: int) =
for i in 0 ..< n:
buf[i and 1023] = $i
The integers go from 0 to 100 million (10 million in Python), so each string has up to eight digits. I report the best of five runs on an Apple M4 Max. C++ is compiled with -O3, Rust in release mode, Nim with -d:danger.

Millions of strings per second:
| million strings per second | |
|---|---|
C++ std::to_string
|
183.8 |
Nim $i
|
85.9 |
Go strconv.Itoa
|
84.3 |
Node.js String(i)
|
73.1 |
Rust itoa
|
71.0 |
Bun String(i)
|
68.5 |
Rust to_string()
|
63.6 |
Python str(i)
|
22.9 |
Python is the slowest at 44 ns per string, about three times slower than JavaScript and eight times slower than C++.
The compiled languages that allocate each string on the heap (Rust, Go, Nim) end up in the same range as JavaScript: 12 to 16 ns per string. Rust with its standard to_string() is even a bit slower than Node.js and Bun. Garbage-collected runtimes like Go and JavaScript are very good at allocating many small, short-lived objects.
C++ wins by a wide margin at 5.4 ns per string. The trick is the small string optimization: a std::string stores short strings, directly inside the object. Our strings have at most eight digits, so C++ never calls the memory allocator.
Versions used: macOS 15.7.7, CPython 3.14.5, Node.js 25.9.0, Bun 1.4.2, Apple clang 17.0.0 (clang-1700.6.4.2) with libc++, rustc 1.94.1 with itoa 1.0.18, Go 1.24.3, Nim 2.2.12.
2026-09-24 22:09:39

Mainstream economists have been advocating for the marvellous effects of global trade for decades. Who needs these factory jobs anyway? We’ll be designing the robots and the nuclear rockets.
Except that, no. It does not work like that.
If you have the factories, sooner or later, you get the designers. If you don’t have the factories, you lose the expertise. Or you never get it.
There aren’t that many people left in the Silicon Valley capable of working on silicon. You can design a chip from California. Making one is done where the fab is.
You can get a PhD in robotics in Quebec City, but you won’t be designing robots unless you fly over to where they are made. Zoom calls won’t cut it at scale.
Jensen Huang’s account, as Joseph Steinberg reported it, is that American manufacturing jobs declined because the work was outsourced. Steinberg says this is flatly wrong, and that technology accounts for the vast majority of the decline in manufacturing’s share of employment.
In 2000 there were 17.3 million manufacturing jobs in the United States. The peak was 19.6 million, in June 1979. From the early 1980s to 2000 the count stayed high, apart from the recessions. China joined the WTO in December 2001. By 2010 manufacturing employment was 11.5 million. In August 2026 it was 12.6 million.
That is 5.8 million jobs gone in a decade. About a million have come back since the bottom. The rest have not.
If technology were the main driver, output should have kept rising while the jobs fell. The fifteen years before 2000 are what that looks like. Manufacturing output nearly doubled, from an index of 51 to an index of 93 (2017 = 100). Employment went from 17.8 million to 17.3 million. The robots were already here in 1985. Employment did not collapse.
After 2000, production stalled. Manufacturing output peaked just under 107 in December 2007. In August 2026 the index was 99.1. Total industrial production was 103.1. American factories are not turning out a flood of extra goods. They are turning out roughly what they turned out twenty years ago.
Output per hour did rise while the jobs were disappearing. The BLS index of manufacturing labor productivity went from about 70 in 2000 to about 100 in 2010. Then it stopped. In early 2026 it was still about 100. Flat output, fewer workers, a higher ratio. The ratio has been flat for fifteen years.
| Year | Jobs (millions) | Manufacturing output (2017 = 100) |
|---|---|---|
| 1985 | 17.8 | 51 |
| 2000 | 17.3 | 93 |
| 2007 | 13.9 | 105 |
| 2010 | 11.5 | 93 |
| August 2026 | 12.6 | 99.1 |
The problem after 2000 was not that American factories became too productive. What changed after 2000 was where the goods were made.
And now, often, we don’t know anymore how to make things. Human expertise matters, and you maintain it by building stuff locally. You are not going to design microprocessors in Maine. It just won’t happen.
2026-09-22 10:27:44

I maintain and comaintain several open-source libraries. Some of them are widely used: ada parses URLs in Node.js, fast_float parses numbers in GCC’s standard library and in Chromium, simdjson parses JSON in Node.js, simdutf validates and transcodes Unicode in Node.js, and the Roaring bitmap libraries sit inside many database engines.
These libraries are mature. They have been optimized for years, by me and by others. For a long time, their performance was flat. Not because nobody cared, but because the remaining gains were expensive: each one required a few days of careful work, and nobody had the days.
Then, in 2026, six of them got much faster, most of it in a few weeks of summer.
To formalize my feeling, I rebuilt every commit of each library from scratch and benchmarked it on one machine (an Intel Xeon Gold 6548N). I track the speedup over time relative to August 2024. Thus the value 1.0 means no speedup. Whereas 2.0 means that the performance doubled. The lines are steps because performance only changes at a commit.
I should say that I cannot know how much AI was involved in each instance. I don’t ask how people arrived at their code. All I ask is that it be good. As for myself, I code with Claude (Opus 5), Grok and DeepSeek (V4 Pro). I was an early adopter of Grok for coding, and it got really good over time.

The roaring library is the Go version of the Roaring index data structure. Decoding to an array got 2.5 times faster, the multi-way union FastOr got 3.1 times faster on one data set, the many-value iterator got 4.5 to 5.9 times faster, and the intersection cardinality gained 10%.
One of the contributors is an AI, actually. It is perfloop. (Disclosure: I am an advisor for perfloop.)
I did a lot of work. We also got help from Philipp Klose who declared using Claude.

The ada library is a standard compliant URL parser. From August 2024 to July 2026, about 550 commits went in and the throughput on a corpus of 100,000 URLs stayed at 0.54 GB/s. Then, in six weeks, it went to 1.28 GB/s: 2.4 times faster, about 15 million URLs per second on one core.
Most of the optimizations were done by Yagiz Nizipli, my long-time co-author. Yagiz works at SpaceX and uses Cursor (presumably with a grok model). Abdul Rawoof Khan and Dillon Mulroy also contributed an optimization each. I worked at optimizing IP address parsing, but it won’t show in this particular benchmark.

The fast_float library parses floating-point numbers from text. It is part of GCC and most browsers. Performance was flat for fifteen months. Then, from March to July 2026, it gained 43% on one file (canada.txt, long coordinates) and 70% on another (mesh.txt, short coordinates). The optimizations should be credited to Koleman Nix and Filipe Oliveira.

The simdjson library recently gained support for C++26 static reflection: you serialize and parse your own structs directly, with no glue code. Since February 2026, serialization is 1.6 times faster on twitter.json and 2.1 times faster on citm_catalog.json. Deserialization, JSON straight into a struct, gained a more modest 10% and 14% (the second panel). (The reflection code only exists since early 2026.) The number of instructions per byte fell by almost exactly the same ratio as the throughput rose: from 6.1 to 3.1 instructions per byte on citm_catalog.json serialization.
Francisco Geiman Thiesen (Microsoft) did most of the work on the serialization side while I mostly helped improve our parsing. Francisco uses Claude.

The simdutf library validates and transcodes UTF-8, UTF-16 and UTF-32, and encodes and decodes base64. ASCII validation went from 83 GB/s to 160 GB/s. UTF-16 validation went from 62 GB/s to 102 GB/s. Base64 decoding gained 17%.
The work was done by Yagiz Nizipli (again) and myself.
The library got other amazing optimizations that do not show up on this benchmark by Gaspard Petit and Shreesh Adiga.

CRoaring implements Roaring bitmaps in C. On the real data sets from the repository, membership tests (contains) got 2.4 times faster, the cardinality of 64-bit bitmaps got 4.9 times faster, iterating over a 64-bit bitmap got 1.9 times faster, decoding a dense bitmap to an array got 2.2 times faster. Unions gained a more modest 13% to 16%.
The authors were Andrei Gudkov and myself.
The techniques used are all well-known. So why all these optimizations all of a sudden? Simply put, in my view, because it got cheap to try new ideas.
There is a lot of talk about the risks of AI in software. Human beings tend to be susceptible to the one-sided bet fallacy: when we see the downsides, we tend to ignore the benefits. Cars kill people, but ambulances save them.
In this instance, the benefits are concrete. Millions of people run these libraries, and this summer, they got faster.