MoreRSS

site iconDaniel LemireModify

Computer science professor at the University of Quebec (TELUQ), open-source hacker, and long-time blogger.
Please copy the RSS to your reader, or quickly subscribe to:

Inoreader Feedly Follow Feedbin Local Reader

Rss preview of Blog of Daniel Lemire

simdjson 5.0 is out

2026-09-28 20:38:25

simdjson 5.0 is out

The simdjson library is a C++ library to parse and generate JSON. It is used in Node.js, ClickHouse, Meta Velox, StarRocks, Apache Doris, the Ladybird browser and many other systems. We released version 4.0 a year ago, in September 2025. Its headline feature was C++26 static reflection: you could turn a C++ structure into JSON, and back, without writing any glue code.

Today we are releasing version 5.0.

  1. Static reflection is no longer guarded and is an officially supported feature. When your compiler has reflection enabled (e.g., g++ -std=c++26 -freflection with GCC 16), simdjson detects it by itself and the reflection-based functions become available.

  2. When deserializing a C++ structure through reflection, simdjson now uses key selectors (see below) by default: it reads the object in a single pass, whatever the order of the keys. You can return to the previous approach (one lookup per member) with -DSIMDJSON_DISABLE_KEY_SELECTOR_REFLECTION=1.

  3. A positive integer in [2^64, 10^20) is now reported as a big integer (BIGINT_NUMBER), like other overflowing integers, instead of a malformed number. When big integers are parsed as strings, a token such as 123456789123456789123x is now rejected.

One major change is the key selectors. A common task is to extract a few fields from a JSON object. In simdjson 5.0 (C++20 or better), you can name the keys at compile time and visit the object once:

using namespace simdjson;
auto json = R"({ "name": "Daniel", "age": 42, "city": "Montreal" })"_padded;
ondemand::parser parser;
auto doc = parser.iterate(json);
std::string_view name, city;
uint64_t age = 0;
auto result = doc.get_object().for_each<"name", "city", "age">(name, city, age);
// name == "Daniel", city == "Montreal", age == 42

The keys can appear in any order. At compile time, simdjson builds a perfect hash function for your set of keys. At run time, recognizing a key takes a hash computed from a couple of bytes and one comparison. You can also pass one callback per key instead of variables. The iteration stops as soon as all keys have been found.

We added annotations for the data structures for automated (C++26) serialization and deserialization: rename, rename_all, alias, skip, default_value, flatten, deny_unknown_fields, transparent, and so forth.

struct [[= simdjson::rename_all<simdjson::case_style::camel_case>]] User {
  std::string first_name;
  int64_t user_id;
  [[= simdjson::rename<"KEY">]] int api_key;
};
// {"firstName":"Ann","userId":7,"KEY":8}

We often get many JSON documents in one file or one network message. simdjson has long supported streams of documents separated by white space (NDJSON). In 5.0 we added:

  • RFC 7464 JSON text sequences (each document preceded by the record separator character) and comma-separated documents;
  • stream_format::newline_delimited: you promise that each document sits on its own line. When you only read part of a document, simdjson jumps to the next line instead of walking over the rest of the document.
  • simdjson::slice_at, which cuts a stream into blocks at document boundaries so you can parse the blocks on as many threads as you like. The built-in threaded mode uses at most two threads.

We also fixed several bugs in document_stream, found in an audit by Francisco Geiman Thiesen.

There are many other smaller features.

  • NaN and infinity: JSON does not allow NaN or Infinity, but many systems produce them anyway. If you define SIMDJSON_ENABLE_NAN_INF, simdjson parses them, and serializes them.
  • Narrow types: get_uint8(), get_int8(), get_uint16(), get_int16() check the range for you. With C++23, get_float32() and get_float64() return std::float32_t and std::float64_t. The binary32 value is rounded once, directly from the decimal string, not through a double.
  • The DOM API can parse a buffer that has no padding (parser.parse_unpadded(...)). It is slower than the regular function, but it never reads past the end of your buffer and it does not copy your data.
  • With C++17, simdjson::padded_input adds padding only when it is needed: when your string ends near a page boundary.
  • C++20 ranges: you can pipe On-Demand arrays and objects into std::views::transform and other adaptors.
  • DOM arrays support reverse iteration (rbegin(), rend()) with no allocation.
  • On-Demand objects offer get_current_position() and revert_position(): if you miss an optional field, you can go back to where you were instead of rescanning the whole object.
  • char8_t (u8) variants of the string accessors in C++20.
  • Better pretty printing with the FracturedJson style, including tables.
  • Memory-mapped files under Windows.
  • We support the memory-safe compiler Fil-C.

The simdjson 5.0 release improved performance compared to simdjson 4.0 in some key cases. Let me review some of them.

I built both versions with GCC 16.1 (-O3, CMake Release) and ran them on an Intel Xeon Gold 6548N (Emerald Rapids), pinned to one core.

Let us start with DOM parsing of our standard files (GB/s):

file 4.0 5.0 speedup
twitter 4.78 4.82 1.0
citm_catalog 4.82 4.79 1.0
github_events 5.33 5.30 1.0
canada 1.10 1.21 1.1
marine_ik 1.25 1.38 1.1
mesh 1.17 1.26 1.1
numbers 1.13 1.41 1.3
twitterescaped 1.59 2.85 1.8
update-center 3.96 3.84 1.0
apache_builds 4.95 4.78 1.0

Files full of numbers (canada, marine_ik, mesh, numbers) are 8% to 25% faster. And a file full of escaped Unicode characters (twitterescaped) is almost twice as fast: among other changes, we now decode consecutive uXXXX sequences without going back to the string scanner between them.

We also serialize faster. Printing floating-point numbers used to be a bottleneck. We replaced the ancient Grisu2 by Dragonbox, and removed calls to memcpy and memmove from the hot path.

file 4.0 5.0 speedup
twitter 0.94 0.96 1.0
citm_catalog 1.07 1.08 1.0
gsoc-2018 1.10 1.25 1.1
canada 0.31 0.52 1.7
marine_ik 0.28 0.37 1.3
mesh 0.34 0.47 1.4
numbers 0.32 0.49 1.6

The simdjson library is a community project. Since version 4.6, contributions came from fior512, 吴杨帆, Alecto Irene Perez, Francisco Geiman Thiesen, Max Bachmann, MoonFlowww, Advit Arora, Jaël Champagne Gareau, Taimoor Kiani, Vasily Pelikh, jmestwa-coder, liyinlong, AlbertoFVisconti, Aylin Dmello, Cuda Chen, Ezra Li, Madhurendra Purbay, Makkar, Pastoray, Paul Dreik, Pavel Kruglov, Piotr Kubaj, Yusuf İhsan Görgel, metsw24-max, neil, pratap singh, Vladimir Saraikin, wankun, xaldarof, Riyane El Qoqui, Justin Li and others. Thank you!

How fast can you fix a UTF-16 string in C#?

2026-09-27 00:53:53

How fast can you fix a UTF-16 string in C#

C# strings are UTF-16. Most characters are one 16-bit code unit. Characters outside the basic multilingual plane, emoji included, take two: a high surrogate (U+D800 to U+DBFF) followed by a low surrogate (U+DC00 to U+DFFF). A surrogate with the wrong neighbor, or with none, is ill-formed.

You should never send an ill-formed string to disk or to the network. It is a bad practice.

In JavaScript, we have fast functions to fix strings or check whether they need fixing:

  • String.prototype.toWellFormed() replaces every lone surrogate with U+FFFD.
  • isWellFormed() reports whether any replacement is needed.

I added both functions to my C# library SimdUnicode, in pull request 54. The algorithm is the same that we contributed to the JavaScript engine V8, so Chrome already fixes strings this way.

string s = UTF16.ToWellFormed(input); // same instance, when the input is already well formed
bool ok = UTF16.IsWellFormed(span);

When the input is well formed, ToWellFormed returns it as is. No allocation.

How are strings fixed? Basically, you replace bad inputs by the replacement character U+FFFD.

Our processors have special instructions called SIMD that allow data parallelism: you can compare multiple values at once. Recent x64 processors from AMD and Intel have better data parallelism than ARM chips, although both have powerful instructions.

The conventional approach in C# to repair a string is a function such as the following.

static void Repair(ReadOnlySpan<char> input, Span<char> output)
{
    input.CopyTo(output);
    int i = NextError(output, 0);
    while (i >= 0)
    {
        output[i] = 'uFFFD';
        i = NextError(output, i + 1);
    }
}
// Index of the next lone surrogate at or after 'start', or -1 if none.
static int NextError(ReadOnlySpan<char> s, int start)
{
    int i = start;
    while (true)
    {
        int k = s.Slice(i).IndexOfAnyInRange('uD800', 'uDFFF');
        if (k < 0) return -1;
        i += k;
        if (char.IsHighSurrogate(s[i]) && i + 1 < s.Length && char.IsLowSurrogate(s[i + 1]))
            i += 2; // valid pair, skip it
        else
            return i;
    }
}

In SimdUnicode, I also use data parallelism.

Let me measure.

Is the UTF-16 string well formed? Intel Xeon Gold 6548N

On the Xeon, with AVX-512, Latin validates at 69 GB/s against 33 GB/s for IndexOfAnyInRange. The Emoji input is well formed, and it is nothing but surrogate pairs. The runtime search drops to 0.4 GB/s. Our check holds 53 GB/s.

Is the UTF-16 string well formed? Apple M4 Max

Our results are similar on the M4 Max, although a bit less impressive compared to the Intel results.

Validation can return at the first lone surrogate. The buffer form of ToWellFormed writes every code unit, a copy of the input or U+FFFD. When the input is well formed, it is effectively a memory copy. Thus we can compare the performance against a copy.

Copy the string, replace lone surrogates. Intel Xeon Gold 6548N

Copy the string, replace lone surrogates. Apple M4 Max

Roughly speaking, we are consistently about as fast as a copy.

Versions used: .NET SDK 10.0.400 on Linux, 10.0.103 on macOS. Intel Xeon Gold 6548N (Emerald Rapids). Apple M4 Max.

Clausecker, R., & Lemire, D. (2026). Fixing ill-formed UTF-16 strings with SIMD instructions. Software: Practice and Experience. (arXiv)

Source code.

A thesis isn’t enough for a PhD

2026-09-26 10:48:13

To get a PhD, you typically have to enroll in a graduate program. Then you complete a few relatively easy courses.
 
You might need to pass a comprehensive examination that checks whether you have a basic understanding of the field. A few students fail at this point, but not many.Then you write a thesis and defend it.
 
In theory, you could write a strong thesis and still fail the oral defense. That is uncommon. The defense is typically a public event, and there may be guests. Failing a student at that stage would be a public humiliation. I have seen students who could not answer basic questions still receive their PhDs because the thesis itself was good enough.
 
Thus, to a very good approximation, getting a PhD amounted to writing a thesis.But we have a problem. AI can write something that looks quite a bit like a thesis. With a bit of prompting, you can produce something that looks like a PhD thesis.I have been telling everyone I can that we have a big problem.Apparently other people have realized this as well. People at Harvard are now saying that a thesis is insufficient for a PhD.
 
 
 
« The dissertation has in the past frequently been used as a proxy for the kind of mathematical development of a student that we expect: acquiring mathematical knowledge and demonstrating independent achievement. However, with modern AI systems, dissertations are not (…) a reliable tool for evaluating students, and PhDs should not be awarded primarily on the basis of the text of the dissertation. We thus recommend regular, multi-faceted evaluations of students in person, by multiple faculty, as the central component of assessment. These evaluations must be rigorous, not pro-forma, and should produce reports that might become part of the student’s dossier for future employment. » 

How many strings can you create per second?

2026-09-25 09:30:49

How many strings can you create per second?

We create new strings all the time. How quickly can you produce short strings in your programming languge? To create a meaningful string, we convert an integer to a string. In Python, that is str(i). The loop stores each new string in a small ring buffer of 1024 slots, so that the engine cannot simply discard the work.

def from_int(n):
    b = buf
    for i in range(n):
        b[i & 1023] = str(i)

In JavaScript (Node.js and Bun), I use String(i):

function from_int(n) {
  for (let i = 0; i < n; i++) buf[i & 1023] = String(i);
}

In C++, I use std::to_string:

void from_int(uint64_t n) {
  for (uint64_t i = 0; i < n; i++) buf[i & 1023] = std::to_string(i);
}

In Rust, I use the standard to_string() and, as an alternative, the popular itoa crate, which is designed for fast integer formatting:

fn std_to_string(buf: &mut [String], n: u64) {
    for i in 0..n {
        buf[(i & 1023) as usize] = i.to_string();
    }
}
fn itoa_to_string(buf: &mut [String], n: u64) {
    let mut b = itoa::Buffer::new();
    for i in 0..n {
        buf[(i & 1023) as usize] = b.format(i).to_owned();
    }
}

In Go, I use strconv.Itoa:

func fromInt(n int) {
    for i := 0; i < n; i++ {
        buf[i&1023] = strconv.Itoa(i)
    }
}

In Nim, I use the $ operator:

proc fromInt(n: int) =
  for i in 0 ..< n:
    buf[i and 1023] = $i

The integers go from 0 to 100 million (10 million in Python), so each string has up to eight digits. I report the best of five runs on an Apple M4 Max. C++ is compiled with -O3, Rust in release mode, Nim with -d:danger.

Time to convert an integer to a new string on an Apple M4 Max

Millions of strings per second:

million strings per second
C++ std::to_string 183.8
Nim $i 85.9
Go strconv.Itoa 84.3
Node.js String(i) 73.1
Rust itoa 71.0
Bun String(i) 68.5
Rust to_string() 63.6
Python str(i) 22.9

Python is the slowest at 44 ns per string, about three times slower than JavaScript and eight times slower than C++.

The compiled languages that allocate each string on the heap (Rust, Go, Nim) end up in the same range as JavaScript: 12 to 16 ns per string. Rust with its standard to_string() is even a bit slower than Node.js and Bun. Garbage-collected runtimes like Go and JavaScript are very good at allocating many small, short-lived objects.

C++ wins by a wide margin at 5.4 ns per string. The trick is the small string optimization: a std::string stores short strings, directly inside the object. Our strings have at most eight digits, so C++ never calls the memory allocator.

Versions used: macOS 15.7.7, CPython 3.14.5, Node.js 25.9.0, Bun 1.4.2, Apple clang 17.0.0 (clang-1700.6.4.2) with libc++, rustc 1.94.1 with itoa 1.0.18, Go 1.24.3, Nim 2.2.12.

Source code.

If you don’t have the factories, you lose the expertise

2026-09-24 22:09:39

Mainstream economists have been advocating for the marvellous effects of global trade for decades. Who needs these factory jobs anyway? We’ll be designing the robots and the nuclear rockets.

Except that, no. It does not work like that.

If you have the factories, sooner or later, you get the designers. If you don’t have the factories, you lose the expertise. Or you never get it.

There aren’t that many people left in the Silicon Valley capable of working on silicon. You can design a chip from California. Making one is done where the fab is.

You can get a PhD in robotics in Quebec City, but you won’t be designing robots unless you fly over to where they are made. Zoom calls won’t cut it at scale.

Jensen Huang’s account, as Joseph Steinberg reported it, is that American manufacturing jobs declined because the work was outsourced. Steinberg says this is flatly wrong, and that technology accounts for the vast majority of the decline in manufacturing’s share of employment.

U.S. manufacturing employment, millions of jobs

In 2000 there were 17.3 million manufacturing jobs in the United States. The peak was 19.6 million, in June 1979. From the early 1980s to 2000 the count stayed high, apart from the recessions. China joined the WTO in December 2001. By 2010 manufacturing employment was 11.5 million. In August 2026 it was 12.6 million.

That is 5.8 million jobs gone in a decade. About a million have come back since the bottom. The rest have not.

If technology were the main driver, output should have kept rising while the jobs fell. The fifteen years before 2000 are what that looks like. Manufacturing output nearly doubled, from an index of 51 to an index of 93 (2017 = 100). Employment went from 17.8 million to 17.3 million. The robots were already here in 1985. Employment did not collapse.

U.S. industrial production, index 2017 = 100

After 2000, production stalled. Manufacturing output peaked just under 107 in December 2007. In August 2026 the index was 99.1. Total industrial production was 103.1. American factories are not turning out a flood of extra goods. They are turning out roughly what they turned out twenty years ago.

Output per hour did rise while the jobs were disappearing. The BLS index of manufacturing labor productivity went from about 70 in 2000 to about 100 in 2010. Then it stopped. In early 2026 it was still about 100. Flat output, fewer workers, a higher ratio. The ratio has been flat for fifteen years.

Year Jobs (millions) Manufacturing output (2017 = 100)
1985 17.8 51
2000 17.3 93
2007 13.9 105
2010 11.5 93
August 2026 12.6 99.1

The problem after 2000 was not that American factories became too productive. What changed after 2000 was where the goods were made.

And now, often, we don’t know anymore how to make things. Human expertise matters, and you maintain it by building stuff locally. You are not going to design microprocessors in Maine. It just won’t happen.

 

A summer of AI optimization

2026-09-22 10:27:44

A summer of AI optimization

I maintain and comaintain several open-source libraries. Some of them are widely used: ada parses URLs in Node.js, fast_float parses numbers in GCC’s standard library and in Chromium, simdjson parses JSON in Node.js, simdutf validates and transcodes Unicode in Node.js, and the Roaring bitmap libraries sit inside many database engines.

These libraries are mature. They have been optimized for years, by me and by others. For a long time, their performance was flat. Not because nobody cared, but because the remaining gains were expensive: each one required a few days of careful work, and nobody had the days.

Then, in 2026, six of them got much faster, most of it in a few weeks of summer.

To formalize my feeling, I rebuilt every commit of each library from scratch and benchmarked it on one machine (an Intel Xeon Gold 6548N). I track the speedup over time relative to August 2024. Thus the value 1.0 means no speedup. Whereas 2.0 means that the performance doubled. The lines are steps because performance only changes at a commit.

I should say that I cannot know how much AI was involved in each instance. I don’t ask how people arrived at their code. All I ask is that it be good. As for myself, I code with Claude (Opus 5), Grok and DeepSeek (V4 Pro). I was an early adopter of Grok for coding, and it got really good over time.

1. roaring (compressed bitmaps, Go)

roaring: speedup over time

The roaring library is the Go version of the Roaring index data structure. Decoding to an array got 2.5 times faster, the multi-way union FastOr got 3.1 times faster on one data set, the many-value iterator got 4.5 to 5.9 times faster, and the intersection cardinality gained 10%.

One of the contributors is an AI, actually. It is perfloop. (Disclosure: I am an advisor for perfloop.)

I did a lot of work. We also got help from Philipp Klose who declared using Claude.

2. ada (URL parsing)

ada: speedup over time

The ada library is a standard compliant URL parser. From August 2024 to July 2026, about 550 commits went in and the throughput on a corpus of 100,000 URLs stayed at 0.54 GB/s. Then, in six weeks, it went to 1.28 GB/s: 2.4 times faster, about 15 million URLs per second on one core.

Most of the optimizations were done by Yagiz Nizipli, my long-time co-author. Yagiz works at SpaceX and uses Cursor (presumably with a grok model). Abdul Rawoof Khan and Dillon Mulroy also contributed an optimization each. I worked at optimizing IP address parsing, but it won’t show in this particular benchmark.

3. fast_float (number parsing)

fast_float: speedup over time

The fast_float library parses floating-point numbers from text. It is part of GCC and most browsers. Performance was flat for fifteen months. Then, from March to July 2026, it gained 43% on one file (canada.txt, long coordinates) and 70% on another (mesh.txt, short coordinates). The optimizations should be credited to Koleman Nix and Filipe Oliveira.

4. simdjson (JSON serialization and deserialization with C++26 reflection)

simdjson: serialization and deserialization speedup over time

The simdjson library recently gained support for C++26 static reflection: you serialize and parse your own structs directly, with no glue code. Since February 2026, serialization is 1.6 times faster on twitter.json and 2.1 times faster on citm_catalog.json. Deserialization, JSON straight into a struct, gained a more modest 10% and 14% (the second panel). (The reflection code only exists since early 2026.) The number of instructions per byte fell by almost exactly the same ratio as the throughput rose: from 6.1 to 3.1 instructions per byte on citm_catalog.json serialization.

Francisco Geiman Thiesen (Microsoft) did most of the work on the serialization side while I mostly helped improve our parsing. Francisco uses Claude.

5. simdutf (Unicode validation and transcoding)

simdutf: speedup over time

The simdutf library validates and transcodes UTF-8, UTF-16 and UTF-32, and encodes and decodes base64. ASCII validation went from 83 GB/s to 160 GB/s. UTF-16 validation went from 62 GB/s to 102 GB/s. Base64 decoding gained 17%.

The work was done by Yagiz Nizipli (again) and myself.

The library got other amazing optimizations that do not show up on this benchmark by Gaspard Petit and Shreesh Adiga.

6. CRoaring (compressed bitmaps, C)

CRoaring: speedup over time

CRoaring implements Roaring bitmaps in C. On the real data sets from the repository, membership tests (contains) got 2.4 times faster, the cardinality of 64-bit bitmaps got 4.9 times faster, iterating over a 64-bit bitmap got 1.9 times faster, decoding a dense bitmap to an array got 2.2 times faster. Unions gained a more modest 13% to 16%.

The authors were Andrei Gudkov and myself.

What happened

The techniques used are all well-known. So why all these optimizations all of a sudden? Simply put, in my view, because it got cheap to try new ideas.

There is a lot of talk about the risks of AI in software. Human beings tend to be susceptible to the one-sided bet fallacy: when we see the downsides, we tend to ignore the benefits. Cars kill people, but ambulances save them.

In this instance, the benefits are concrete. Millions of people run these libraries, and this summer, they got faster.