Skip to main content

Module pp

Module pp 

Source
Expand description

The C99 preprocessor: translation phase 4.

The preprocessor sits between the lexer and the parser. It consumes the lexer’s tokens together with their bol / preceded_by_space flags, executes the directives it finds and replaces macro invocations, and hands the parser a Token list that contains no # directives at all.

§Token origin

Every token the preprocessor emits carries an Origin:

  • Origin::Source — the token was written where it is, and its Token::range is its own.
  • Origin::Expansion — the token came out of a macro’s replacement list, and its range is the range of the invocation, not of the #define.

That distinction is the whole point. A diagnostic — ours or rustc’s on the code we generate — must land on something the user wrote, and the #define is not where the mistake is being made. Tokens that came from a macro argument keep their own ranges, because the argument was written at the call site; only the replacement list’s own tokens, and the tokens # and ## synthesise, are re-pointed at the invocation.

A procedural macro cannot emit secondary spans, so the “which macro was that?” half of the story is appended to the message instead: Expansions::annotate adds note: in expansion of macro 'X' to every diagnostic that lands inside an invocation.

§Hide sets

Macro replacement follows Dave Prosser’s algorithm, the one the standard’s rescanning rules were written from. Each token carries a hide set: the names of the macros whose expansion it came out of. A name is not replaced again while it is in its own hide set, which is what stops

#define foo (4 + foo)

from running forever while still letting foo be replaced somewhere else. For a function-like macro the hide set of the result is (HS(name) ∩ HS(')')) ∪ {name}, which is what makes the standard’s f(2 * (f)(z)) example come out right.

§The # # rule

Rust’s own lexer refuses ## in raw-token mode (“reserved multi-hash token”), so a c99! block written as raw Rust tokens cannot spell the token-pasting operator. It can spell a # # b, and in a replacement list a # immediately followed by another # is ill-formed C anyway — # must be followed by a macro parameter — so this preprocessor reads two adjacent # tokens in a replacement list as the ## operator. The rule applies in every input mode, so a macro written with # # means the same thing whether it is passed as raw tokens or inside a string literal.

§#include

A header is read at the point the directive is reached, lexed, and pushed onto a stack of open files; the tokens it produces are the tokens the parser sees next. Each file has its own text, its own line numbering and its own idea of what __FILE__ says, and each is placed in a range of the global offset space of its own, so that a position identifies both a file and a place in it. The Preprocessed::included list is what a caller adds to its source map afterwards, in the order the files were opened; the preprocessor cannot do it itself, because it runs on a thread where a proc_macro2::Span cannot follow it.

Where a header is looked for — and why the system directories are never looked in — is crate::include. Reading one twice is avoided the two usual ways: #pragma once, and the classic include-guard optimisation. A file that includes itself with neither eventually nests too deeply and is reported.

A conditional group a header leaves open ends with the header rather than running on into whatever included it, and is reported against the file that opened it.

§The cinrs pragmas

#pragma cinrs target "i686-unknown-linux-gnu"
#pragma cinrs include_path "vendor/include"
#pragma cinrs system_include first
#pragma cinrs link "mylib"
#pragma cinrs export
#pragma cinrs safe gcd fact
#pragma cinrs no_std
#pragma cinrs crate "crate::vendor::cinrs"

target picks the data model the unit is translated for, overriding CINRS_TARGET; include_path adds a directory to the search path (relative paths resolve against CARGO_MANIFEST_DIR); system_include puts the platform’s own include directories on that path, after the bundled headers or — with first — before them (see crate::include); link adds #[link(name = "mylib")] on an extern block of its own; export gives everything with external linkage a real C symbol, so that another unit can link to it; safe generates those functions without unsafe, so that rustc checks them (see crate::sema::check_safe); no_std takes the Vecs a variable length array or alloca needs from alloc rather than from std; and crate says where the cinrs facade crate is, for the generated code that names the runtime. Being directives rather than attributes or macro arguments is what makes them mean the same thing in raw-token and in string-literal input. An unknown #pragma cinrs option is an error; every other pragma is ignored, as 6.10.6 asks. doc/pragmas.md is the reference page.

target is the one that cannot be handled where it stands: the predefined macros are built from the model before the first directive is read, so scan_target_pragma finds it lexically, before preprocessing, and the handler here only checks that what it finds agrees. Which is why the pragma has to be written in the unit’s own text, ahead of any #include or #if; anywhere else is a diagnostic rather than a silent half-measure.

§What the later revisions add

__VA_OPT__(…), #elifdef and #elifndef are C23’s, and are accepted in a Standard::C23 block; an older one is told which macro would have them. true and false are keywords there too, so an #if reads them as 1 and 0 rather than turning them into 0 like any other identifier (C23 6.10.1p6). #embed and __has_embed are C23’s as well: the directive is replaced by the bytes of a file, written as a comma-separated list of unsigned char values, and the resource is reported in Preprocessed::embedded_files so that editing it rebuilds the crate.

§Predefined macros

MacroValue
__STDC__1
__STDC_HOSTED__1
__STDC_VERSION__the revision: 199901L, 201112L, 201710L or 202311L; undefined in c89! and gnu89!
__cinrs__, __CINRS__1
__CINRS_MAJOR__, __CINRS_MINOR__, __CINRS_PATCH__the version of cinrs-core
__GNUC__, __GNUC_MINOR__, __GNUC_PATCHLEVEL__14, 2, 0
__VERSION__"14.2.0 (cinrs <version>)"
__FILE__the invoking .rs file’s path, or "<c99!>"
__LINE__the line of the invoking .rs file
__DATE__"??? ?? ????"
__TIME__"??:??:??"

__FILE__ and __LINE__ are computed from the position they are used at, so inside a header they name the header and the line in it, and a macro defined in <assert.h> that mentions them reports the line the assertion is written on.

__DATE__ and __TIME__ are deliberately fixed placeholders: a build has to be reproducible, and a macro that expanded to the wall clock would make the generated code differ between two builds of the same source.

__LINE__ is a line of the .rs file the invocation is written in whenever the compiler tells us where that is — the captured C text remembers which line of its .rs file it starts on — and a line inside the C text itself otherwise (a TokenStream built from a string in a unit test, for instance).

§#line

#line N and #line N "name" (6.10.4), and GCC’s # N "name" flags… line marker, do what they say: the line after the directive is line N, counting up per physical line from there, and __FILE__ is the given name until the next directive or the end of that file. The macro-expanded form is supported too — #line line, with line a macro — and the numbering is per file, so a #line inside a header ends with the header. In the macro’s own text a #line replaces the .rs-line convention above from the next line to the end of the block, which is exactly what a program that writes one is asking for.

Nothing else moves. A diagnostic — this crate’s or rustc’s — still points at the token that was really written, in the file it was really written in, because that is the position the user can look at; making the caret land on the C is the reason the whole pipeline carries spans. __BASE_FILE__ names the file the translation unit started in and is not affected either; __FILE_NAME__ is __FILE__ without the directory, so it is.

On top of those comes a small, deliberately short set of target description macros derived from the machine this crate was compiled for and from TargetModel: the architecture (__x86_64__, __aarch64__, …), the operating system (__linux__, __unix__, _WIN32, __APPLE__, …), the data model (__LP64__, __ILP32__, __CHAR_UNSIGNED__, __SIZEOF_INT__ and friends, __CHAR_BIT__) and the byte order (__BYTE_ORDER__).

As a compiler, cinrs says it is GCC 14.2 — __GNUC__ is what the world’s version gates test, and 4.2.1 turned real programs away or onto their slow paths — and says who it really is with __CINRS__. The GNU macros a header tests for are defined where cinrs does what they promise (__GNUC_STDC_INLINE__, __GCC_HAVE_SYNC_COMPARE_AND_SWAP_n, the __GCC_ATOMIC_* family, __BIGGEST_ALIGNMENT__), __GCC_IEC_559 and __GCC_IEC_559_COMPLEX are 0 because Annexes F and G are not claimed, and the rest are left out on purpose: __OPTIMIZE__ and __NO_INLINE__, __GCC_ASM_FLAG_OUTPUTS__ (flag outputs are refused), __SIZEOF_FLOAT128__, __PRAGMA_REDEFINE_EXTNAME and __clang__.

Structs§

Context
What the preprocessor needs to know about the text it is running over.
Expansion
One macro expansion a token came out of.
Expansions
Every macro invocation the preprocessor replaced, so that a diagnostic landing inside one can say which macro it was.
IncludedFile
One file #include brought in, for the caller to add to its source map.
PackMap
What #pragma pack asked for, at every point of the token list.
Preprocessed
Everything one run of the preprocessor produced.
SafeName
One function #pragma cinrs safe named.
TargetOptionMap
What #pragma GCC target asked for, at every point of the token list.
TargetPragmas
What scan_target_pragma found, which the preprocessor needs in order not to report the same directive twice.
Token
A preprocessed C token: what the parser consumes.

Enums§

Origin
Where a preprocessed token came from.

Constants§

DEFAULT_FILE_NAME
What __FILE__ expands to when the compiler will not say where the invocation is.

Functions§

preprocess
Runs the preprocessor over a lexed translation unit.
scan_target_pragma
Finds #pragma cinrs target "<triple>" in a freshly lexed unit and puts the model it names into options.