Expand description
The C99 preprocessor: translation phase 4.
The preprocessor sits between the lexer and the
parser. It consumes the lexer’s tokens together with their
bol / preceded_by_space flags, executes the directives it finds and
replaces macro invocations, and hands the parser a Token list that
contains no # directives at all.
§Token origin
Every token the preprocessor emits carries an Origin:
Origin::Source— the token was written where it is, and itsToken::rangeis its own.Origin::Expansion— the token came out of a macro’s replacement list, and its range is the range of the invocation, not of the#define.
That distinction is the whole point. A diagnostic — ours or rustc’s on
the code we generate — must land on something the user wrote, and the
#define is not where the mistake is being made. Tokens that came from a
macro argument keep their own ranges, because the argument was written
at the call site; only the replacement list’s own tokens, and the tokens
# and ## synthesise, are re-pointed at the invocation.
A procedural macro cannot emit secondary spans, so the “which macro was
that?” half of the story is appended to the message instead:
Expansions::annotate adds note: in expansion of macro 'X' to every
diagnostic that lands inside an invocation.
§Hide sets
Macro replacement follows Dave Prosser’s algorithm, the one the standard’s rescanning rules were written from. Each token carries a hide set: the names of the macros whose expansion it came out of. A name is not replaced again while it is in its own hide set, which is what stops
#define foo (4 + foo)from running forever while still letting foo be replaced somewhere else.
For a function-like macro the hide set of the result is
(HS(name) ∩ HS(')')) ∪ {name}, which is what makes the standard’s
f(2 * (f)(z)) example come out right.
§The # # rule
Rust’s own lexer refuses ## in raw-token mode (“reserved multi-hash
token”), so a c99! block written as raw Rust tokens cannot spell the
token-pasting operator. It can spell a # # b, and in a replacement list
a # immediately followed by another # is ill-formed C anyway — # must
be followed by a macro parameter — so this preprocessor reads two adjacent
# tokens in a replacement list as the ## operator. The rule applies in
every input mode, so a macro written with # # means the same thing whether
it is passed as raw tokens or inside a string literal.
§#include
A header is read at the point the directive is reached, lexed, and pushed
onto a stack of open files; the tokens it produces are the tokens the
parser sees next. Each file has its own text, its own line numbering and
its own idea of what __FILE__ says, and each is placed in a range of the
global offset space of its own, so that a position identifies both a file
and a place in it. The Preprocessed::included list is what a caller
adds to its source map afterwards, in the order the
files were opened; the preprocessor cannot do it itself, because it runs on
a thread where a proc_macro2::Span cannot follow it.
Where a header is looked for — and why the system directories are never
looked in — is crate::include. Reading one twice is avoided the two
usual ways: #pragma once, and the classic include-guard optimisation. A
file that includes itself with neither eventually nests too deeply and is
reported.
A conditional group a header leaves open ends with the header rather than running on into whatever included it, and is reported against the file that opened it.
§The cinrs pragmas
#pragma cinrs target "i686-unknown-linux-gnu"
#pragma cinrs include_path "vendor/include"
#pragma cinrs system_include first
#pragma cinrs link "mylib"
#pragma cinrs export
#pragma cinrs safe gcd fact
#pragma cinrs no_std
#pragma cinrs crate "crate::vendor::cinrs"target picks the data model the unit is translated for, overriding
CINRS_TARGET; include_path adds a directory to the search path (relative
paths resolve against CARGO_MANIFEST_DIR); system_include puts the
platform’s own include directories on that path, after the bundled headers
or — with first — before them (see crate::include); link adds
#[link(name = "mylib")] on an extern block of its own; export gives
everything with external linkage a real C symbol, so that another unit can
link to it; safe generates those functions without unsafe, so that
rustc checks them (see crate::sema::check_safe); no_std takes the
Vecs a variable length array or alloca needs from alloc rather than
from std; and crate says where the cinrs facade crate is, for the
generated code that names the runtime. Being
directives rather than attributes or macro arguments is what makes them
mean the same thing in raw-token and in string-literal input. An unknown
#pragma cinrs option is an error; every other pragma is ignored, as
6.10.6 asks. doc/pragmas.md is the reference page.
target is the one that cannot be handled where it stands: the predefined
macros are built from the model before the first directive is read, so
scan_target_pragma finds it lexically, before preprocessing, and the
handler here only checks that what it finds agrees. Which is why the pragma
has to be written in the unit’s own text, ahead of any #include or #if;
anywhere else is a diagnostic rather than a silent half-measure.
§What the later revisions add
__VA_OPT__(…), #elifdef and #elifndef are C23’s, and are accepted in
a Standard::C23 block; an older one is told which macro would have
them. true and false are keywords there too, so an #if reads them as
1 and 0 rather than turning them into 0 like any other identifier (C23
6.10.1p6). #embed and __has_embed are C23’s as well: the directive is
replaced by the bytes of a file, written as a comma-separated list of
unsigned char values, and the resource is reported in
Preprocessed::embedded_files so that editing it rebuilds the crate.
§Predefined macros
| Macro | Value |
|---|---|
__STDC__ | 1 |
__STDC_HOSTED__ | 1 |
__STDC_VERSION__ | the revision: 199901L, 201112L, 201710L or 202311L; undefined in c89! and gnu89! |
__cinrs__, __CINRS__ | 1 |
__CINRS_MAJOR__, __CINRS_MINOR__, __CINRS_PATCH__ | the version of cinrs-core |
__GNUC__, __GNUC_MINOR__, __GNUC_PATCHLEVEL__ | 14, 2, 0 |
__VERSION__ | "14.2.0 (cinrs <version>)" |
__FILE__ | the invoking .rs file’s path, or "<c99!>" |
__LINE__ | the line of the invoking .rs file |
__DATE__ | "??? ?? ????" |
__TIME__ | "??:??:??" |
__FILE__ and __LINE__ are computed from the position they are used
at, so inside a header they name the header and the line in it, and a macro
defined in <assert.h> that mentions them reports the line the assertion
is written on.
__DATE__ and __TIME__ are deliberately fixed placeholders: a build has
to be reproducible, and a macro that expanded to the wall clock would make
the generated code differ between two builds of the same source.
__LINE__ is a line of the .rs file the invocation is written in
whenever the compiler tells us where that is — the captured C text
remembers which line of its .rs file it starts on — and a line inside the
C text itself otherwise (a TokenStream built from a string in a unit
test, for instance).
§#line
#line N and #line N "name" (6.10.4), and GCC’s # N "name" flags… line
marker, do what they say: the line after the directive is line N, counting
up per physical line from there, and __FILE__ is the given name until the
next directive or the end of that file. The macro-expanded form is
supported too — #line line, with line a macro — and the numbering is
per file, so a #line inside a header ends with the header. In the
macro’s own text a #line replaces the .rs-line convention above from
the next line to the end of the block, which is exactly what a program that
writes one is asking for.
Nothing else moves. A diagnostic — this crate’s or rustc’s — still
points at the token that was really written, in the file it was really
written in, because that is the position the user can look at; making the
caret land on the C is the reason the whole pipeline carries spans.
__BASE_FILE__ names the file the translation unit started in and is not
affected either; __FILE_NAME__ is __FILE__ without the directory, so it
is.
On top of those comes a small, deliberately short set of target
description macros derived from the machine this crate was compiled for and
from TargetModel: the architecture
(__x86_64__, __aarch64__, …), the operating system (__linux__,
__unix__, _WIN32, __APPLE__, …), the data model (__LP64__,
__ILP32__, __CHAR_UNSIGNED__, __SIZEOF_INT__ and friends,
__CHAR_BIT__) and the byte order (__BYTE_ORDER__).
As a compiler, cinrs says it is GCC 14.2 — __GNUC__ is what the world’s
version gates test, and 4.2.1 turned real programs away or onto their slow
paths — and says who it really is with __CINRS__. The GNU macros a header
tests for are defined where cinrs does what they promise
(__GNUC_STDC_INLINE__, __GCC_HAVE_SYNC_COMPARE_AND_SWAP_n, the
__GCC_ATOMIC_* family, __BIGGEST_ALIGNMENT__), __GCC_IEC_559 and
__GCC_IEC_559_COMPLEX are 0 because Annexes F and G are not claimed, and
the rest are left out on purpose: __OPTIMIZE__ and __NO_INLINE__,
__GCC_ASM_FLAG_OUTPUTS__ (flag outputs are refused),
__SIZEOF_FLOAT128__, __PRAGMA_REDEFINE_EXTNAME and __clang__.
Structs§
- Context
- What the preprocessor needs to know about the text it is running over.
- Expansion
- One macro expansion a token came out of.
- Expansions
- Every macro invocation the preprocessor replaced, so that a diagnostic landing inside one can say which macro it was.
- Included
File - One file
#includebrought in, for the caller to add to its source map. - PackMap
- What
#pragma packasked for, at every point of the token list. - Preprocessed
- Everything one run of the preprocessor produced.
- Safe
Name - One function
#pragma cinrs safenamed. - Target
Option Map - What
#pragma GCC targetasked for, at every point of the token list. - Target
Pragmas - What
scan_target_pragmafound, which the preprocessor needs in order not to report the same directive twice. - Token
- A preprocessed C token: what the parser consumes.
Enums§
- Origin
- Where a preprocessed token came from.
Constants§
- DEFAULT_
FILE_ NAME - What
__FILE__expands to when the compiler will not say where the invocation is.
Functions§
- preprocess
- Runs the preprocessor over a lexed translation unit.
- scan_
target_ pragma - Finds
#pragma cinrs target "<triple>"in a freshly lexed unit and puts the model it names intooptions.