rucc_codegen/lower.rs
1//! The selector: an IR function becomes a machine IR function.
2//!
3//! Design: `spec/10-backend.md` sections 10.2 and 10.3.
4//!
5//! What the matcher in [`crate::select`] does is answer one question about one term. What this
6//! does is ask it: walk a function, decide which terms are worth asking about, and build machine
7//! instructions out of what comes back. Nothing here decides what an IR term lowers to. That is
8//! in `rules/x86-64.rules` and it is proved before it is used, which is the whole point of the
9//! arrangement and the reason this file is short.
10//!
11//! # What it does with an instruction
12//!
13//! It tries the ways the instruction can be shown to the matcher, in order, and takes the first
14//! that a rule fires on. [`crate::term`] is what a way of showing one is, and the order is the
15//! most specific first: an operand that is a constant is offered as a constant before it is
16//! offered as a register, and an operand computed by an instruction of its own is offered as
17//! that instruction before it is offered as a register. A rule that wants an immediate too wide
18//! for the machine has a guard that turns it down, and the search carries on to the way of
19//! showing it that puts the constant in a register, which is the right answer and is one nobody
20//! had to write down.
21//!
22//! A constant is not lowered where it is written. It is materialized where a register for it is
23//! first wanted, which is what keeps a constant that every use folded into an immediate from
24//! leaving a dead instruction behind, and it also gives the value the shortest live range it
25//! could have. The instruction that materializes it comes from the rule set like everything else.
26//!
27//! # What it does not do yet
28//!
29//! Everything is in the general purpose registers, because every rule in the set is about an
30//! integer, so a call that passes a `double` and a function that returns one are both reported
31//! rather than lowered. So is an argument that travels on the stack, on either side of a call,
32//! and so is a call through an address rather than to a name.
33//!
34//! # A call
35//!
36//! Not a rule, because a rule pattern sees one term and what a call's operands are is whatever
37//! the signature made them. [`crate::abi`] builds one instead, out of the same description of the
38//! convention the arguments come from: the values it passes are reads constrained to the
39//! registers the convention places them in, what comes back is a write constrained to the
40//! register it comes back in, and every other register the callee is free to destroy is a write
41//! of that register and nothing else, which is all the allocator needs to keep a value out of it.
42//!
43//! What that costs the frame is an argument area, and nothing after selection could work out how
44//! big, so the size of the widest call is given back with the function. A function that makes no
45//! call at all is a leaf, and a leaf is the function that may use the red zone.
46//!
47//! # Where a block goes
48//!
49//! On the block, which is what machine IR does with an edge and is why the branches need no more
50//! rule language than the arithmetic did. A rule never names a block, so an unconditional jump
51//! has no rule at all and a conditional branch has one that is about its condition and nothing
52//! else. The arms are copied across after the block is filled, arguments and all, because an
53//! argument that is a constant is materialized where a register for it is first wanted and the
54//! end of the block is where an edge wants it.
55//!
56//! What this leaves behind is a function whose blocks are in the order the IR held them and whose
57//! branches are still branches on a register. Turning one into a `test` and a `jcc` is the block
58//! layout's, since which of the two arms falls through is the layout's answer, and [`crate::split`]
59//! has to run before allocation so that every edge carrying a value has somewhere to put it.
60//!
61//! A store and a return are the two things here that write no register. A store is emitted like
62//! everything else and the only difference is that there is no result to put anywhere, so the
63//! operands the target describes are all reads. A return is the same, and what it is for is its
64//! one operand: the target constrains it to the register the caller reads the value out of, and
65//! the allocator is what gets it there. The instruction that leaves is not chosen here at all,
66//! because the epilogue has to give the frame back first and [`crate::finish`] writes that after
67//! allocation, so a return of nothing is lowered to nothing.
68//!
69//! The entry block is the one block whose parameters are not block parameters here. They are the
70//! function's arguments, they are already somewhere when it starts, and [`crate::abi`] is what
71//! says where. An argument that arrives on the stack is reported rather than read, because where
72//! the stack put it is a distance into a frame and no frame exists until after allocation.
73//!
74//! Blocks are walked in the order the function holds them and a value is expected to be defined
75//! before it is used, which is true of the IR this is given because every pass before it keeps
76//! definitions ahead of uses.
77
78use std::collections::HashSet;
79use std::fmt;
80
81use rucc_base::{Interner, Symbol};
82use rucc_diag::Span;
83use rucc_ir::{
84 Abi, AsmOperand, AsmOperands, AttrSet, Block, Def, Extra, Flags, FloatPred, Func, Inst,
85 Linkage, MemOrder, Opcode, Param, PrefetchHint, RmwOp, Type, Value, Visibility,
86};
87use rucc_mir as mir;
88use rucc_target::x86_64;
89use rucc_target::{CallRegs, Constraint, OperandDesc, PhysReg, RegClass, Role, Segment};
90
91use crate::abi::{self, Missing, Refused};
92use crate::coverage::Fired;
93use crate::elsewhere::Elsewhere;
94use crate::frame::{Layout, Local};
95use crate::select::{Match, Piece, Rule, Table};
96use crate::term::{MAX_ARGS, PLAIN, Plan, Shown, Term, Terms};
97use crate::varargs;
98
99/// The prefix a rule file puts in front of a machine term, which says which target it belongs
100/// to and is not part of the opcode.
101pub(crate) const PREFIX: &str = "x64.";
102
103/// The instruction a global offset table slot is read with.
104///
105/// Not in [`x86_64::FRAME`] with the other opcodes this file names, because a frame has no use for
106/// it. It is spelled out here because the relocation it takes is only legal on a `mov` with a REX
107/// prefix, so the width is part of the requirement rather than a choice.
108const GOT_LOAD: &str = "mov_rm_64";
109
110/// The instruction a template's `jmp` to a name outside it becomes.
111///
112/// Named here for [`GOT_LOAD`]'s reason turned round: a frame never writes one, because the only
113/// function it appears in has no prologue and no epilogue for the frame to write anything into.
114/// See [`x86_64::Step::Away`].
115const AWAY: &str = "jmp_away";
116
117/// How wide an address is on this target, which is the width a cast between a pointer and an
118/// integer has to be at for the cast to be nothing.
119const ADDRESS_BITS: u32 = 64;
120
121/// How much of a register an operand of an `asm` statement fills, which is the width of its type
122/// with two exceptions. A pointer is an address, and a truth value is the byte it is stored in: a
123/// program that writes `sete %0` into a `_Bool` is asking for exactly that byte, which is what tcc's
124/// own test of the width of one checks.
125fn held_bits(ty: Type) -> u32 {
126 if ty.is_ptr() {
127 ADDRESS_BITS
128 } else if ty.bits() == 1 {
129 8
130 } else {
131 ty.bits()
132 }
133}
134
135/// How many bytes a `long double` takes in memory, and what it is aligned to, which are the same
136/// number and are both more than the ten bytes that mean anything.
137///
138/// The psABI's answer rather than a choice here. `sizeof (long double)` is sixteen on this
139/// machine, so an array of them is laid out this way whatever a slot holding one does, and a slot
140/// that agreed with the array is one fewer thing to get wrong.
141const X87_BYTES: u32 = 16;
142
143/// How many values the x87 stack holds at once.
144///
145/// Eight, which is the machine's number rather than a choice here, and it matters in one place:
146/// the parameters of a block are copied through the stack so that they all move at once, and a
147/// block with more of them than this has nowhere to put the ninth.
148const X87_DEPTH: usize = 8;
149
150/// How far into the buffer of a `__builtin_setjmp` each of the four words it writes is.
151///
152/// The first three are gcc's, measured against gcc 16.2.0 on x86-64 at `-O0`: the frame pointer,
153/// the address control comes back to, and the stack pointer, in that order. The fourth is this
154/// compiler's own. gcc has no word for the answer because it writes a second block that sets the
155/// answer to one and is arrived at from the restore, and this writes the answer through memory
156/// instead, for the reason [`Lowering::saves_place`] gives.
157///
158/// None of the four is an interface. The buffer is the program's memory and its five words are
159/// the front end's promise about how much of it there is, but nothing except the matching restore
160/// ever reads a word of it, and a buffer written by one compiler was never going to be one another
161/// compiler could come back through.
162const JUMP_FRAME: i32 = 0;
163
164/// Where the address control comes back to is. See [`JUMP_FRAME`].
165const JUMP_PC: i32 = 8;
166
167/// Where the stack pointer is. See [`JUMP_FRAME`].
168const JUMP_STACK: i32 = 16;
169
170/// Where the address of the word the answer arrives in is. See [`JUMP_FRAME`].
171const JUMP_ANSWER: i32 = 24;
172
173/// How many bytes the word a `__builtin_setjmp` answers with takes in the frame, and what it is
174/// aligned to, which are the same number because it is one machine word.
175const JUMP_WORD: u32 = 8;
176
177/// How many registers the restore needs to hold things in while it puts the frame back.
178///
179/// Four, and every one of them is a register nothing else in the function may be in, which is why
180/// they are counted here rather than asked for one at a time. See [`Lowering::comes_back`].
181const JUMP_REGS: usize = 4;
182
183/// How many bytes a value passes through on its way between a register and the x87 stack.
184///
185/// Eight, because the widest thing that crosses is a `double` or a sixty four bit integer, and
186/// nothing crosses at eighty bits: a value that wide is already in the frame and the stack reaches
187/// it where it is.
188const X87_CROSSING: u32 = 8;
189
190/// Where the rounding field of the x87 control word is and what it has to be set to for the unit
191/// to cut towards zero, which is the one rounding C asks for that the unit does not do by default.
192///
193/// Both bits on is truncate. The field is ORed into the word that was already there rather than
194/// written over it, so the precision control and the exception masks somebody else set stay set.
195const X87_TRUNCATE: i64 = 0x0c00;
196
197/// Whether a type is the one this machine has no register for.
198///
199/// Only the eighty bit float is, and that is a fact about x86-64 rather than about floats: every
200/// other scalar the front end produces is in a general purpose register or a vector one, and this
201/// one is on the x87 stack while it is being worked on and in memory the rest of the time. So it
202/// has no place in [`Lowering::class_of`] and no name in [`crate::term`], and every instruction
203/// that touches one is written out by hand in this file.
204fn on_x87(ty: Type) -> bool {
205 ty.is_scalar() && ty.is_float() && ty.bits() == 80
206}
207
208/// Where one operand of an assembly statement is, on each side of the assembly.
209///
210/// Two registers rather than one, because an operand written `+` is a value that arrives and a
211/// value that leaves and those are two values. The machine IR has one definition per register by
212/// construction, so an instruction of the template that reads the operand and writes it has to name
213/// a different register in each place, and what makes the two one register in the end is the
214/// [`Constraint::Reuse`] the instruction's description carries: the allocator reads it, gives both
215/// the same physical register, and copies the incoming value somewhere first when something else is
216/// still using it.
217///
218/// Most operands have one of the two. An input has only a place it is read from and an output
219/// written `=` has only a place it is written to, and asking either of them for the other is an
220/// operand read where the opcode writes or written where it reads, which [`Lowering::placed`]
221/// refuses.
222#[derive(Debug, Clone, Copy, Default, PartialEq, Eq)]
223struct Place {
224 /// The register the value arrives in, for an operand something reads.
225 read: Option<mir::Reg>,
226 /// The register the value leaves in, for an operand something writes.
227 write: Option<mir::Reg>,
228}
229
230/// Whether that operand of the statement is one the assembly may read, and so where a read of it
231/// gets its value from.
232///
233/// [`bound`] asks this question of an operand a constraint letter named and this asks it of one the
234/// template numbered, which is the same question twice because a two-address instruction reaches
235/// its first source both ways. `mulq %3` reaches `rax` by the letter on the output and libgmp says
236/// what is in it with `"%0"` on an input. `addq %5,%q1` reaches its first source by numbering the
237/// output, and libgmp says what is in it with `"0"` on an input in the same way.
238///
239/// So an output written `=` has no value of its own and is still readable when an input is tied to
240/// it, and the value the read wants is that input's. An output written `+` carries its own value
241/// and answers with that. An output nothing is tied to answers `None`, which is a program that told
242/// the compiler the assembly only writes the operand while the instruction reads it before it
243/// writes it, and is refused where it is asked.
244fn read_as(list: &[AsmOperand<'_>], index: usize) -> Option<Value> {
245 let operand = list.get(index)?;
246 if operand.value.is_some() {
247 return operand.value;
248 }
249 operand.result?;
250 list.iter().find(|entry| entry.tied == Some(index)).and_then(|entry| entry.value)
251}
252
253/// Which of an assembly statement's operands is in that register, for an instruction that reaches
254/// the register without its text saying so.
255///
256/// The constraint is what says so, and it is the only thing in such a statement that could:
257/// `"=a"` is an output in `rax`, `"c"` is an input in `rcx`, an operand that is a local register
258/// variable is in the register its declaration named, and a register nothing names is a register
259/// nobody has said anything about. So a write looks among the outputs and a read among the inputs,
260/// and an output written `+` answers for either, since it is read before it is written. See
261/// [`pinned`], which is the one question asked of both ways of saying it.
262///
263/// The other way a read of such a register is said is a matching constraint. `"=a"` on an output
264/// and `"0"` on an input is the program saying that one register holds the input on the way in and
265/// the output on the way out, and it is how a statement fills a register the instruction reads and
266/// writes without writing the register down twice. The letter is on the output, which has no value
267/// to read, and the value is on the input, which has no letter, and the answer is the output: its
268/// place is read out of the register the input arrived in, and in a template with a loop in it the
269/// place moves on to wherever the last write left it, which is what a read on the next time round
270/// wants. tcc steps a pointer along a string with `lodsb` and `"=&S"` tied to `"0"`, and a read of
271/// the input would start the string again every time round.
272///
273/// And a read of a register an output alone is in is a read of that output, the same as a read of
274/// an output the template numbered. tcc copies a string with `lodsb` and `stosb` and `"=&a"` on an
275/// output nothing is tied to, and what `stosb` stores is what `lodsb` loaded one line up, which is
276/// the output as the template left it rather than anything the statement handed in.
277///
278/// `None` is a register the instruction uses and the statement put nothing in, which is the usual
279/// answer rather than an unusual one. `cpuid` writes four registers and a program that wanted one
280/// of them names one. See [`Lowering::spare`], which is where that one goes.
281fn bound(list: &[AsmOperand<'_>], reg: PhysReg, role: Role) -> Option<usize> {
282 let output =
283 list.iter().position(|operand| operand.result.is_some() && pinned(operand) == Some(reg));
284 if role.is_def() {
285 return output;
286 }
287 // The output first when something is in it on the way in, which is what `+` and a matching
288 // constraint both say, since its place is where a write earlier in the template left it and
289 // the read wants that. See [`read_as`] for what it holds before anything wrote it.
290 let arrives = |at: usize| read_as(list, at).is_some();
291 if let Some(at) = output.filter(|&at| arrives(at)) {
292 return Some(at);
293 }
294 let named = list.iter().position(|operand| {
295 operand.result.is_none() && operand.value.is_some() && pinned(operand) == Some(reg)
296 });
297 named.or(output)
298}
299
300/// The register one of an assembly statement's operands is in, whichever of the two ways said it.
301///
302/// A constraint letter is one way and is the only way a program can say one of the six registers
303/// that have a letter. A local register variable is the other, and it is the only way to say any
304/// of the rest: there is no letter for `r12`, which is the whole reason the extension exists, so
305/// the declaration says it and the front end wrote the name into the constraint. The name is read
306/// against this machine's table here, the same place the letter is read against it, and a name the
307/// machine has not got answers nothing, which leaves the operand where an operand nobody placed
308/// goes.
309///
310/// The sigil gcc allows in front of a name is taken off here, because what a name is written with
311/// is syntax and which register it means is this question.
312fn pinned(operand: &AsmOperand<'_>) -> Option<PhysReg> {
313 match operand.named {
314 Some(name) => {
315 let (reg, _) = x86_64::gpr_named(name.strip_prefix('%').unwrap_or(name))?;
316 Some(reg)
317 }
318 None => operand.fixed.and_then(x86_64::gpr_letter),
319 }
320}
321
322/// Why a function could not be lowered.
323///
324/// One reason and then nothing. A function with no rule for something in it is a function this
325/// cannot finish, and the second thing it could not lower is not news.
326#[derive(Debug, Clone, PartialEq, Eq)]
327pub enum Unsupported {
328 /// An instruction no rule fires on.
329 Inst {
330 /// The instruction that stopped it.
331 inst: Inst,
332 /// What the rule file would call it, or nothing if the rule language has no name for it
333 /// at all, which is what an instruction at a width nothing is written about looks like.
334 term: Option<&'static str>,
335 /// The opcode, which is what gets named when the rule language has no word for it.
336 ///
337 /// An opcode the rule language has no word for is exactly the opcode no rule lowers, so
338 /// without this the message would be empty in every case where somebody needs it.
339 opcode: Opcode,
340 /// What it produces, or nothing for an instruction that is only an effect.
341 ty: Option<Type>,
342 },
343 /// A parameter that does not arrive somewhere this can bring it in from.
344 ///
345 /// Not an instruction, which is why it is a separate arm: it is a fact about the signature
346 /// and there is nothing in the body of the function to point at.
347 Argument {
348 /// Its position in the signature.
349 index: usize,
350 /// What is wrong with where it arrives.
351 missing: Missing,
352 },
353 /// A call that passes or gives back a value this cannot put where the convention wants it.
354 Call {
355 /// The call.
356 inst: Inst,
357 /// Which value, and what is wrong with where it travels.
358 refused: Refused,
359 },
360 /// A `return` this cannot put where the convention wants it.
361 ///
362 /// A separate arm from [`Unsupported::Inst`] because it is not an instruction no rule fires
363 /// on. A return of more than one value is built from the convention rather than matched, the
364 /// same way a call is, so what goes wrong with one is what goes wrong with a call and not the
365 /// absence of a rule.
366 Returned {
367 /// The `return`.
368 inst: Inst,
369 /// What is wrong with where one of the values travels.
370 missing: Missing,
371 },
372 /// A stack slot the frame cannot give the bytes it asked for.
373 ///
374 /// Not an instruction no rule covers. An `alloca` is built here rather than matched, so what
375 /// goes wrong with one is what the frame can and cannot hold rather than what the rules spell.
376 Dynamic {
377 /// The `alloca`.
378 inst: Inst,
379 /// What the frame could not do about it.
380 growing: Growing,
381 },
382 /// More parameters of a type that travels on the x87 stack than the stack is deep.
383 ///
384 /// Not an instruction either, for the reason a function's parameter is not one: it is a fact
385 /// about the block and there is nothing in the block to point at. What crosses an edge for one
386 /// of these is the address of where the value is, and the block copies the bytes into a slot
387 /// of its own, all of them through the stack at once so that a block carrying two of them
388 /// swapped is copied in an order that is right. Eight is as many as the stack holds, and a
389 /// ninth would have to be copied before or after the rest, which is the order that could be
390 /// wrong.
391 Phi {
392 /// Which block it arrives at.
393 block: Block,
394 /// How many of them arrive there, which is the whole of what is wrong.
395 count: usize,
396 /// What they are.
397 ty: Type,
398 },
399 /// An `asm` statement this cannot build.
400 ///
401 /// Not an instruction no rule fires on, for the reason a call is not one: what it stands for is
402 /// whatever its template says, and no pattern over terms can read a string.
403 Assembly {
404 /// The `inline_asm`.
405 inst: Inst,
406 /// What about it is not built here yet.
407 refused: Written,
408 },
409 /// A `register long x asm ("...")` naming something this machine has not got.
410 ///
411 /// Not an instruction no rule fires on. There is a rule's worth of instruction here and what
412 /// is wrong is the string beside it, which is a name rather than a term, so the message says
413 /// the name. Which names a machine has is the machine's own question and this is where it is
414 /// asked, at the table a clobber list is read against.
415 Register {
416 /// The `register_value`.
417 inst: Inst,
418 /// The name the program wrote, as it wrote it.
419 name: String,
420 },
421 /// A naked function whose frame is not empty.
422 ///
423 /// Not an instruction no rule fires on, and there is nothing in the body to point at: the
424 /// function asked for no prologue and then wanted bytes only a prologue takes. Refused rather
425 /// than given the bytes anyway, because an offset into a frame nothing set up reaches into
426 /// whatever the caller left below its own stack pointer, which is wrong code that assembles.
427 /// See [`crate::frame::Layout::naked`].
428 Naked {
429 /// How many bytes it wanted, which is the whole of what is wrong.
430 bytes: u32,
431 },
432}
433
434/// What about an `asm` statement is not built yet.
435#[derive(Debug, Clone, Copy, PartialEq, Eq)]
436pub enum Written {
437 /// A template with instructions in it.
438 Template,
439 /// An `asm goto`, whose labels make the statement a terminator.
440 Goto,
441 /// An operand this cannot put where the constraint says it goes.
442 Operand,
443 /// A clobber list naming something this has no register for.
444 Clobber,
445 /// A `jmp` out of the function in a function that has an epilogue behind it.
446 Away,
447}
448
449impl Written {
450 /// The rest of the sentence that starts with the statement.
451 #[must_use]
452 pub fn why(self) -> &'static str {
453 match self {
454 // The template is the assembler's to read and there is no assembler here yet, so a
455 // template with anything in it is a string nothing can turn into bytes. An empty one is
456 // no instructions, and no instructions is something this can write.
457 Written::Template => "has instructions in its template, which nothing here assembles",
458 Written::Goto => "jumps to a label, which nothing here builds an edge for",
459 Written::Operand => "has an operand this cannot place",
460 Written::Clobber => "says it destroys a register this has no name for",
461 Written::Away => {
462 "jumps out of the function, which only a function that is `naked` may do, since \
463 anywhere else there is an epilogue behind it to give the frame back"
464 }
465 }
466 }
467}
468
469/// What the frame could not do about a stack slot.
470#[derive(Debug, Clone, Copy, PartialEq, Eq)]
471pub enum Growing {
472 /// An object of a size the number a frame counts bytes in does not reach.
473 Huge,
474 /// A variable length array wanting more alignment than a call leaves the stack pointer with.
475 ///
476 /// Rounding the stack pointer down again after the bytes have been taken would put it
477 /// somewhere no constant reaches the rest of the frame from, so a frame like this needs a
478 /// second base register held for the whole of the function. Nothing here holds one.
479 ///
480 /// [`crate::expand::rounds`] takes the array away before this sees it, by asking for the
481 /// alignment in extra bytes and handing out an address inside them, so what is left of this
482 /// is IR that arrived without going through that pass and the fixed local in
483 /// [`crate::pipeline`] that wants the same thing from the other side.
484 Aligned,
485 /// A variable length array in a function written without a prologue.
486 ///
487 /// A frame that grows is reached from a frame pointer, and establishing one is the first two
488 /// instructions of a prologue that `__attribute__((naked))` asked there be none of. See
489 /// [`crate::frame::Layout::naked`].
490 Naked,
491}
492
493impl Growing {
494 /// The rest of the sentence that starts with the slot.
495 #[must_use]
496 pub fn why(self) -> &'static str {
497 match self {
498 Growing::Huge => "is more bytes than a frame counts",
499 Growing::Aligned => {
500 "wants more alignment than the stack pointer is left on, which needs a base \
501 register nothing here keeps"
502 }
503 Growing::Naked => {
504 "is in a function that is `naked`, which has no prologue to point a frame pointer \
505 at it with"
506 }
507 }
508 }
509}
510
511impl Unsupported {
512 /// The instruction it is about, or nothing for the one arm that is about a signature.
513 ///
514 /// What a caller wants this for is the span. The function knows where every instruction in
515 /// it came from, so a caller holding both can point a message at the line somebody wrote
516 /// rather than at the file as a whole, and nothing here has to carry a span of its own.
517 pub fn inst(&self) -> Option<Inst> {
518 match *self {
519 Unsupported::Inst { inst, .. }
520 | Unsupported::Call { inst, .. }
521 | Unsupported::Returned { inst, .. }
522 | Unsupported::Dynamic { inst, .. }
523 | Unsupported::Assembly { inst, .. }
524 | Unsupported::Register { inst, .. } => Some(inst),
525 Unsupported::Argument { .. } | Unsupported::Phi { .. } | Unsupported::Naked { .. } => {
526 None
527 }
528 }
529 }
530}
531
532impl fmt::Display for Unsupported {
533 fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
534 match *self {
535 Unsupported::Inst { term: Some(term), .. } => write!(f, "no rule lowers `{term}`"),
536 Unsupported::Inst { term: None, opcode, ty: Some(ty), .. } => {
537 write!(f, "no rule lowers a `{opcode}` producing a `{ty}`")
538 }
539 Unsupported::Inst { term: None, opcode, ty: None, .. } => {
540 write!(f, "no rule lowers a `{opcode}`")
541 }
542 Unsupported::Argument { index, missing } => {
543 write!(f, "parameter {index} {}", missing.why())
544 }
545 Unsupported::Call { refused: Refused { argument: Some(index), missing }, .. } => {
546 write!(f, "argument {index} of this call {}", missing.why())
547 }
548 Unsupported::Call { refused: Refused { argument: None, missing }, .. } => {
549 write!(f, "what this call gives back {}", missing.why())
550 }
551 Unsupported::Returned { missing, .. } => {
552 write!(f, "what this function gives back {}", missing.why())
553 }
554 Unsupported::Dynamic { growing, .. } => {
555 write!(f, "this local {}", growing.why())
556 }
557 Unsupported::Phi { block, count, ty } => {
558 let block = block.index();
559 write!(
560 f,
561 "block{block} takes {count} parameters of type `{ty}` and only {X87_DEPTH} can cross an edge at once"
562 )
563 }
564 Unsupported::Assembly { refused, .. } => write!(f, "this `asm` {}", refused.why()),
565 Unsupported::Register { ref name, .. } => {
566 write!(
567 f,
568 "this object is kept in `{name}`, which is not a register this machine has"
569 )
570 }
571 Unsupported::Naked { bytes } => write!(
572 f,
573 "this function is `naked` and wants {bytes} bytes of frame, which there is no prologue to take"
574 ),
575 }
576 }
577}
578
579impl std::error::Error for Unsupported {}
580
581/// A lowered function, and what the frame needs that the machine IR does not hold.
582#[derive(Debug)]
583pub struct Lowered {
584 /// The function, in machine instructions.
585 pub func: mir::Func,
586 /// What it wants its stack to look like, which is separate from the function so that the two
587 /// can be read and written at the same time.
588 pub stack: Stack,
589 /// Which rules of the table lowered it, which is what `-Zrule-coverage` asks for and what
590 /// `crate::coverage` writes down.
591 pub fired: Fired,
592 /// Which machine IR block each IR block became, indexed by the IR block's own index, and
593 /// nothing for a block the walk never reached.
594 ///
595 /// Here because it is the only place the correspondence exists. Selection makes one block per
596 /// block, in the same order and with the arms in the same order, so anything the IR knows
597 /// about a block can be carried down through this and nothing else, and
598 /// [`crate::weights::carry`] is what does.
599 pub blocks: Vec<Option<mir::Block>>,
600}
601
602/// What a function's stack has to hold, as far as selection is able to say.
603///
604/// All of it is answered here because selection is where a call is built and where an `alloca`
605/// is read, and nothing after it could tell what either of them needed.
606#[derive(Debug, Default)]
607pub struct Stack {
608 /// How many bytes the widest call in the function needs below the stack pointer for the
609 /// arguments it passes there, or `None` for a function that makes no call at all.
610 ///
611 /// `None` is a leaf, which is the function that may use the red zone and the one whose stack
612 /// pointer does not have to be left aligned for anybody.
613 pub calls: Option<u32>,
614 /// The memory the function asked for itself, one entry for every `alloca` in it, in the order
615 /// the walk reached them.
616 pub locals: Vec<Local>,
617 /// Which instruction computes the address of which of those locals.
618 ///
619 /// An address in the frame is a distance from the stack pointer, and there is no frame until
620 /// after allocation, so the instruction is written here with nothing in its displacement and
621 /// [`crate::finish`] writes the number in once [`crate::frame::Frame`] knows it.
622 pub addresses: Vec<(mir::Inst, usize)>,
623 /// Which of those locals is which declaration in the source, for the ones the program declared.
624 ///
625 /// The number is the one the IR function carries and means nothing here. What it is for is the
626 /// debugging information, which has to say where a named local ended up and cannot ask the
627 /// frame directly: the frame knows a local by the order the `alloca` for it was lowered in and
628 /// by nothing else.
629 ///
630 /// Shorter than the list above rather than the same length, because most of what a function
631 /// keeps in its frame is memory an expression wanted somewhere to put.
632 pub declared: Vec<(usize, u32)>,
633 /// Which instruction computes the address of a piece of memory whose size the function works
634 /// out while it runs, which is what a variable length array is.
635 ///
636 /// Waiting on [`crate::finish`] for a different number from the one the addresses above are:
637 /// the bytes were taken off the stack pointer by the instruction in front of this one, so where
638 /// they start is however much of the bottom of the frame belongs to the arguments of a call,
639 /// and that is not known until the frame is.
640 pub dynamic: Vec<mir::Inst>,
641 /// Which instruction takes those bytes off the stack pointer, one for every one of them, in the
642 /// order the walk reached them.
643 ///
644 /// Read by [`crate::finish`] on a command line that asked for the stack to be touched a page at
645 /// a time, which is the one thing that has to find these again: the bytes are in a register by
646 /// then, so the walk down to them is a loop, and a loop is written around an instruction rather
647 /// than in front of a block. Nothing else looks at them, because everything else about a frame
648 /// that grows is answered by the address the instruction below this one computes.
649 pub grown: Vec<mir::Inst>,
650 /// Where the function first moves the stack pointer while it runs, if it does at all.
651 ///
652 /// Two things are read off this. One is whether at all, which is what [`crate::frame::Layout`]
653 /// wants, because a frame that moves its stack pointer has a different shape from one that does
654 /// not and the layout is built before the instructions are looked at again. See `Growing` in
655 /// [`crate::frame`]. The other is where, so that a caller that cannot accept such a frame has
656 /// somewhere to point when it says so.
657 pub grown_at: Option<Inst>,
658 /// Which instruction reads which of the arguments the caller passed on the stack, as how far up
659 /// the caller's argument area it reads.
660 ///
661 /// Waiting on [`crate::finish`] for the same reason the addresses above are, and on one thing
662 /// more: where the caller's argument area is from inside this function depends on whether the
663 /// prologue had to force the stack pointer's alignment, so which register the load reads
664 /// through is not settled here either.
665 pub arguments: Vec<(mir::Inst, u32)>,
666 /// Whether the function asked where its own frame is, which is what `__builtin_frame_address`
667 /// and `__builtin_return_address` both start from.
668 ///
669 /// A function like that keeps a frame pointer whatever the flags say, because the register is
670 /// the answer to the first of them and the start of the walk for every depth above zero. There
671 /// is no other way to reach it: the distance from the stack pointer to the frame is a number
672 /// the layout works out, and what a walk up the chain needs is the link the prologue saved.
673 pub walks_frames: bool,
674 /// Whether the function saved a place for a `__builtin_longjmp` to come back to, which is what
675 /// `__builtin_setjmp` does.
676 ///
677 /// A function like that keeps a frame pointer whatever the flags say as well, and for a reason
678 /// of the same shape: the two registers the restore puts back are the frame pointer and the
679 /// stack pointer, and a frame that did not keep the first of them has nothing in it saying
680 /// where the caller's frame is for the epilogue to find after control has come back.
681 pub saves_place: bool,
682}
683
684impl Stack {
685 /// The layout given, with the three fields only the lowering knows the answer to filled in.
686 ///
687 /// Everything else in a layout comes from the flags the function is compiled under or from the
688 /// allocation, so this takes one and returns it rather than building one.
689 ///
690 /// A function that saved a place is not a leaf whatever it called. What a leaf buys is the red
691 /// zone, which is the words below the stack pointer nothing else may write, and a function
692 /// control comes back into from a `__builtin_longjmp` has already had something else running
693 /// down there: whatever it called and whatever that called, or a signal handler on the same
694 /// stack. Every one of those has written over the red zone by the time control arrives, so a
695 /// value this function left there would not be there any more.
696 #[must_use]
697 pub fn layout<'a>(&'a self, base: Layout<'a>) -> Layout<'a> {
698 Layout {
699 leaf: self.calls.is_none() && !self.saves_place,
700 outgoing: self.calls.unwrap_or(0),
701 locals: &self.locals,
702 grows: self.grown_at.is_some(),
703 ..base
704 }
705 }
706}
707
708/// The x86-64 machine IR for that function.
709///
710/// # Errors
711///
712/// The first instruction no rule fires on, which today is anything at a width the rule set is not
713/// written at, a parameter that does not arrive in a register this can read, or a call that
714/// passes something this cannot put where the convention wants it.
715pub fn func(
716 source: &Func,
717 names: &mut Interner,
718 conv: &'static CallRegs,
719 elsewhere: &Elsewhere,
720) -> Result<Lowered, Unsupported> {
721 Lowering::new(source, names, conv, elsewhere).run()
722}
723
724/// What the matcher settled on for one block, indexed the way the block's instructions are.
725struct Decided {
726 /// What each instruction matched, and nothing for one that matched no rule or was folded
727 /// into a later one.
728 found: Vec<Option<Match<Term>>>,
729 /// How each instruction showed its operands to the matcher, which is what says what it took.
730 plans: Vec<Option<Plan>>,
731 /// The instructions some other instruction took, which are the ones with nothing to write.
732 folded: Vec<Inst>,
733}
734
735/// One function being lowered.
736struct Lowering<'a> {
737 source: &'a Func,
738 names: &'a mut Interner,
739 out: mir::Func,
740 /// The machine register each IR value is in, once it has one.
741 regs: Vec<Option<mir::Reg>>,
742 /// For a constant that has been written into a register, the block it was written into,
743 /// which is the only block that register is any good in.
744 written: Vec<Option<mir::Block>>,
745 /// How many times each IR value is read, which is what says whether an instruction may be
746 /// folded into the one that reads it.
747 uses: Vec<u32>,
748 /// The block being filled.
749 at: Option<mir::Block>,
750 /// The machine IR block each IR block became.
751 blocks: Vec<Option<mir::Block>>,
752 /// The class an address is in, which is the general purpose one and is not a question: every
753 /// register an addressing mode names holds part of an address, and there is no machine here
754 /// that computes an address anywhere but in this file. Which class a *value* is in is
755 /// [`Lowering::class_of`], and it is a question, because a float is in the other one.
756 gpr: RegClass,
757 /// Where the convention this function is compiled for puts things, which is read for the
758 /// arguments and for the calls.
759 conv: &'static CallRegs,
760 /// Which names this function may not work an address out for itself, which is a fact about the
761 /// module and so is worked out before any of this and handed in.
762 elsewhere: &'a Elsewhere,
763 /// What the function wants its stack to look like, filled in as the walk finds out.
764 stack: Stack,
765 /// What a `va_start` in this function has to write, or nothing for a function that takes no
766 /// arguments its signature does not name.
767 ///
768 /// Worked out once, when the entry block binds the parameters, because every number in it is
769 /// about where those parameters left the walk over the argument registers and there is nowhere
770 /// else that knows.
771 varargs: Option<Varargs>,
772 /// Which of the function's stack objects each eighty bit value lives in, once it has asked
773 /// for one.
774 ///
775 /// One slot per value and it is never given back, which is what makes an eighty bit value
776 /// behave like every other one: it is written once and read wherever it is read, and no two
777 /// of them share a slot the way two of them would share a register. What is in a register is
778 /// the address, and that is worked out again at every use rather than kept, so nothing here
779 /// holds a general purpose register open across a whole function.
780 slots: Vec<Option<usize>>,
781 /// The eight bytes a value passes through between a register and the x87 stack, once
782 /// something has wanted them.
783 ///
784 /// One for the whole function, because every group that uses it is a handful of instructions
785 /// with nothing in between: the bytes are written, read straight back and never looked at
786 /// again, so a second slot would be a second slot holding the same nothing.
787 crossing: Option<usize>,
788 /// The four bytes the control word is saved in and the changed copy written to, once
789 /// something has wanted them.
790 ///
791 /// One for the whole function for the reason above, and four rather than two because it is
792 /// two words: the one the unit had and the one with the rounding field turned to truncate.
793 control: Option<usize>,
794 /// The word a `__builtin_setjmp` in this function answers with, once one has asked for it.
795 ///
796 /// One for the whole function however many saves there are in it, because the word is written
797 /// and read back with nothing in between: the save writes a zero into it and the instruction
798 /// straight after reads it, and the only other thing that ever writes it is a restore arriving
799 /// between those two. Two saves sharing it is two pairs each doing that, and neither can be
800 /// inside the other.
801 answer: Option<usize>,
802 /// Which rules have fired so far.
803 fired: Fired,
804}
805
806/// What a `va_start` in a variadic function writes into the list it is given.
807///
808/// Two shapes, because two conventions describe a list two ways, and [`crate::varargs`] is where
809/// both are written down. Neither is a set of numbers on its own: where the save area is and where
810/// the caller's argument area is are distances into a frame that does not exist until after
811/// allocation, so each is a `lea` [`crate::finish`] fills in.
812#[derive(Debug, Clone, Copy, PartialEq, Eq)]
813enum Varargs {
814 /// The four field list, whose two offsets are settled here and whose two addresses are not.
815 Fields {
816 /// Which of the function's stack objects is the register save area.
817 save: usize,
818 /// How far up the caller's argument area the first argument the signature does not name is,
819 /// which is the whole of that area the named ones did not take.
820 incoming: u32,
821 /// What `gp_offset` starts at, which is past the general purpose registers the named
822 /// arguments took.
823 integers: u32,
824 /// What `fp_offset` starts at, which is past the vector ones.
825 floats: u32,
826 },
827 /// The list that is a pointer, which is the one address and nothing else.
828 Pointer {
829 /// How far up the caller's argument area the first argument the signature does not name is,
830 /// which on this convention is the word belonging to the position the named ones stopped
831 /// at.
832 incoming: u32,
833 },
834}
835
836/// How far a function's name reaches, narrowed from the linkage the IR gave it.
837///
838/// The IR has five and an object file says three, and the two the linker cannot tell apart are
839/// the two weak ones: which of them a symbol had is a fact the optimizer reads and the linker has
840/// no way to record. A function is never `Common`, since that is what a tentative definition of an
841/// object is and there is no tentative definition of a function, and it is written here rather
842/// than left out so that a linkage added later has to come past this.
843const fn binding(linkage: Linkage) -> mir::Binding {
844 match linkage {
845 Linkage::Internal => mir::Binding::Local,
846 Linkage::Weak | Linkage::LinkOnce => mir::Binding::Weak,
847 Linkage::External | Linkage::Common => mir::Binding::Global,
848 }
849}
850
851/// How far a function's name reaches outside a shared library, carried across unchanged.
852///
853/// Nothing is narrowed here the way [`binding`] narrows the linkage, because ELF records all
854/// three of these and the two enumerations are the same three answers written twice: once in a
855/// crate that is not allowed to know what an object file is and once in one that is.
856const fn visibility(visibility: Visibility) -> mir::Visibility {
857 match visibility {
858 Visibility::Default => mir::Visibility::Default,
859 Visibility::Hidden => mir::Visibility::Hidden,
860 Visibility::Protected => mir::Visibility::Protected,
861 }
862}
863
864impl<'a> Lowering<'a> {
865 fn new(
866 source: &'a Func,
867 names: &'a mut Interner,
868 conv: &'static CallRegs,
869 elsewhere: &'a Elsewhere,
870 ) -> Self {
871 let counts = source.counts();
872 let name = source.name;
873 let mut uses = vec![0; counts.values];
874 for block in source.blocks() {
875 for inst in source.insts(block) {
876 for &arg in &source[source[inst].args] {
877 uses[arg.index()] += 1;
878 }
879 for call in source.successors(inst) {
880 for &arg in &source[call.args] {
881 uses[arg.index()] += 1;
882 }
883 }
884 }
885 }
886 let mut out = mir::Func::new(name);
887 out.align = source.align;
888 // Carried rather than worked out here, because where a function was declared is a fact
889 // about the source and this is a long way past it. What wants it is the line table.
890 out.declared = source.declared;
891 out.binding = binding(source.linkage);
892 out.visibility = visibility(source.visibility);
893 Self {
894 source,
895 names,
896 out,
897 regs: vec![None; counts.values],
898 written: vec![None; counts.values],
899 blocks: vec![None; counts.blocks],
900 uses,
901 at: None,
902 gpr: x86_64::GPR,
903 conv,
904 elsewhere,
905 stack: Stack::default(),
906 varargs: None,
907 slots: vec![None; counts.values],
908 crossing: None,
909 control: None,
910 answer: None,
911 fired: Fired::new(),
912 }
913 }
914
915 fn run(mut self) -> Result<Lowered, Unsupported> {
916 // Every block before any of them is filled, because a block that jumps forward has to
917 // name the block it jumps to and a machine IR block is named by a handle rather than by
918 // the IR block it came from.
919 for block in self.source.blocks() {
920 let out = self.out.create_block();
921 self.blocks[block.index()] = Some(out);
922 }
923 for block in self.order() {
924 self.block(block)?;
925 }
926 // And the name each block an image holds the address of was given, which nothing in the
927 // walk above would ask for: the `lea` a label address is inside the function needs no
928 // symbol, and the one thing that does is a relocation in another section.
929 let named: Vec<(Block, Symbol)> = self.source.named_blocks().collect();
930 let labels: Vec<(mir::Block, Symbol)> =
931 named.into_iter().map(|(block, name)| (self.out_block(block), name)).collect();
932 self.out.labels = labels;
933 self.naming();
934 Ok(Lowered { func: self.out, stack: self.stack, fired: self.fired, blocks: self.blocks })
935 }
936
937 /// Which register each declaration the front end kept in a value ended up in, as far as this
938 /// walk can say, which is the other half of what [`Lowering::new_reg`] writes down as it goes.
939 ///
940 /// Two halves because there are two ways a value gets a register here. Most of them ask for a
941 /// fresh one and that is where `new_reg` catches them, and the rest are put in a register
942 /// something else chose: a parameter arrives in whichever one the convention handed it, a block
943 /// parameter in whichever one the edge agreed on, and a result of a rule that names its own
944 /// registers in the one the rule named. None of those goes past the mint, so this is the map at
945 /// the end read off the other side, and the two together are every value a declaration is
946 /// behind.
947 ///
948 /// The map on its own would not do, which is why `new_reg` writes down what it writes down: the
949 /// entry for a constant is cleared every time the walk leaves the block that wrote it, so a
950 /// local a constant holds is in the map for one block of the function and nowhere else.
951 fn naming(&mut self) {
952 let mut named = std::mem::take(&mut self.out.named);
953 for value in self.source.values() {
954 let Some(reg) = self.regs[value.index()] else { continue };
955 named.extend(self.source.value_decls(value).map(|decl| (decl, reg)));
956 }
957 named.sort_unstable();
958 named.dedup();
959 self.out.named = named;
960 }
961
962 /// The order the blocks are filled in, which is not the order they are written in.
963 ///
964 /// Reverse postorder, because a value is written in a block that dominates every block that
965 /// reads it and a block in reverse postorder comes before every block it dominates. The order
966 /// the blocks are written in does not have that property: a block written early can read a
967 /// value a block below it writes, and reading a value with no register yet mints one, so the
968 /// register the definition writes later is not the register the read named. Nothing writes the
969 /// one the read named, and what comes out is a function that loads a stack slot no store ever
970 /// reached. It is the order this walk goes in rather than the order the blocks come out in,
971 /// which is what the loop above fixes, so the machine function is still written the way the IR
972 /// function was.
973 ///
974 /// Blocks the entry does not reach come last, in the order they are written in. Nothing runs
975 /// them and nothing they name is read by anything that does, but they still have to be filled,
976 /// because a machine block with no terminator is not one the passes below can read.
977 fn order(&self) -> Vec<Block> {
978 let Some(entry) = self.source.entry() else { return self.source.blocks().collect() };
979 let count = self.blocks.len();
980 let mut succs: Vec<Vec<Block>> = vec![Vec::new(); count];
981 for block in self.source.blocks() {
982 let Some(term) = self.source.terminator(block) else { continue };
983 succs[block.index()] = self.source.successors(term).map(|call| call.block).collect();
984 }
985 // An explicit stack, because the depth of the walk is the number of blocks and a function
986 // built by a generator has as many of those as it likes.
987 let mut seen = vec![false; count];
988 let mut order = Vec::with_capacity(count);
989 let mut stack = vec![(entry, 0usize)];
990 seen[entry.index()] = true;
991 while let Some((block, at)) = stack.pop() {
992 let Some(&next) = succs[block.index()].get(at) else {
993 order.push(block);
994 continue;
995 };
996 stack.push((block, at + 1));
997 if !seen[next.index()] {
998 seen[next.index()] = true;
999 stack.push((next, 0));
1000 }
1001 }
1002 order.reverse();
1003 order.extend(self.source.blocks().filter(|block| !seen[block.index()]));
1004 order
1005 }
1006
1007 /// One block: its parameters, then every instruction in it that is not folded into another.
1008 fn block(&mut self, block: Block) -> Result<(), Unsupported> {
1009 let out = self.out_block(block);
1010 self.at = Some(out);
1011 if self.source.entry() == Some(block) {
1012 self.arrive(block, out)?;
1013 } else {
1014 let mut arriving = Vec::new();
1015 for ¶m in &self.source[block].params {
1016 // A value with no register to arrive in, which the class would not say, since
1017 // `class_of` puts one of these in the general purpose file on purpose and what it
1018 // means by that is that nothing there can hold it. What crosses the edge for one
1019 // of those is the address of where the value already is, so the parameter is a
1020 // pointer here and the bytes it points at are copied below.
1021 let ty = self.source[param].ty;
1022 let reg = self.out.append_param(out, self.class_of(ty));
1023 self.regs[param.index()] = Some(reg);
1024 if on_x87(ty) {
1025 arriving.push((param, reg));
1026 }
1027 }
1028 self.settle(block, &arriving)?;
1029 }
1030
1031 // What each instruction matched, and which instructions were folded into another. The
1032 // decision is made for the whole block before any of it is written, and it is made more
1033 // than once: a value that only some of its readers took has to be put back in a register
1034 // for all of them, and taking it away from those readers changes what they match.
1035 let insts: Vec<Inst> = self.source.insts(block).collect();
1036 let mut refused: HashSet<Value> = HashSet::new();
1037 let mut decided = self.decide(&insts, &refused);
1038 while let Some(value) = self.left_alive(&insts, &decided.plans) {
1039 refused.insert(value);
1040 decided = self.decide(&insts, &refused);
1041 }
1042 let Decided { found, folded, .. } = decided;
1043
1044 for (&inst, matched) in insts.iter().zip(found) {
1045 if folded.contains(&inst) || self.writes_nothing(inst) {
1046 continue;
1047 }
1048 // A call is built from the convention rather than matched, which is why it is the one
1049 // opcode looked at by name here. Through an address it is a different instruction and
1050 // the same convention, so the two arrive at the same place and differ in one line of
1051 // it.
1052 match self.source[inst].opcode {
1053 Opcode::Call | Opcode::CallIndirect => {
1054 self.called(inst)?;
1055 continue;
1056 }
1057 // Built from the frame rather than matched, for the same shape of reason a call
1058 // is built from the convention: what a rule replaces a term with is instructions,
1059 // and what an `alloca` needs first is bytes, which the rule language has no way
1060 // to ask for.
1061 Opcode::Alloca => {
1062 self.reserve(inst)?;
1063 continue;
1064 }
1065 // Reading the stack pointer and writing it back, which are the two ends of a scope
1066 // holding a variable length array. Built here for the reason an `alloca` is: the
1067 // value is a register the rule language has no way to name, because what it holds
1068 // is not a value the program computed but where the machine's stack had got to.
1069 Opcode::StackSave => {
1070 self.stack_pointer(inst, false)?;
1071 continue;
1072 }
1073 Opcode::StackRestore => {
1074 self.stack_pointer(inst, true)?;
1075 continue;
1076 }
1077 // The address of a name, built here for the same reason an `alloca` is: what a
1078 // rule replaces a term with is instructions over values, and the operand of this
1079 // one is a symbol, which is a thing the rule language has no way to bind and the
1080 // solver has no way to say anything about. There is nothing in `lea sym(%rip)` a
1081 // proof over bitvectors could discharge, because what makes it the right answer
1082 // is the relocation and what the linker does with it.
1083 Opcode::GlobalAddr => {
1084 self.address_of(inst)?;
1085 continue;
1086 }
1087 // The address of a label and the branch that reads one, built here for the same
1088 // reason and for one more. The reason is the same: what the first of them names is
1089 // a block, which is not a value a rule pattern can bind, and there is nothing in
1090 // the distance between two places in one function that a proof over bitvectors
1091 // could discharge. The extra one is that the second is a terminator whose arms are
1092 // not two and not fixed, and a rule says what an instruction reads rather than
1093 // where a block goes.
1094 Opcode::BlockAddr => {
1095 self.block_address(inst)?;
1096 continue;
1097 }
1098 Opcode::IndirectBr => {
1099 self.indirect_branch(inst)?;
1100 continue;
1101 }
1102 // A `switch` that `crate::switch` found dense enough for a table, which is a load
1103 // out of the table and the same jump. Built here for the reasons the jump above
1104 // is, and because what the load reads is a place in this function.
1105 Opcode::Switch => {
1106 self.jump_table(inst)?;
1107 continue;
1108 }
1109 // The pair that saves a place in this function and comes back to it. Built here
1110 // for the reason the address of a label is, and for two more. The reason is the
1111 // same: the first of them writes down where control comes back to, which is a
1112 // place in this function and not a value a rule pattern can bind. The extra ones
1113 // are that each of them is a group of instructions over a buffer the program owns
1114 // rather than one instruction, and that the first of them leaves the block it was
1115 // written in and carries on in a new one, which is a thing no rule can do.
1116 Opcode::SetjmpMarker => {
1117 self.saves_place(inst)?;
1118 continue;
1119 }
1120 Opcode::LongjmpMarker => {
1121 self.comes_back(inst)?;
1122 continue;
1123 }
1124 // Where this thread's own storage starts, built here for a reason of the same
1125 // shape: what it reads is `%fs`, which is not a register the rule language can
1126 // bind and not one a proof over bitvectors could say anything about, because what
1127 // makes the load the right answer is an agreement between the loader and the C
1128 // library rather than any arithmetic.
1129 Opcode::ThreadPointer => {
1130 self.thread_pointer(inst)?;
1131 continue;
1132 }
1133 // What a named machine register holds, built here for the reason above written
1134 // about any register rather than about one: which register it is is a string
1135 // beside the instruction, and a rule matches on an opcode and a type and could
1136 // not see it. There is nothing to prove either, since the answer is the register
1137 // and the instruction is the move that reads it.
1138 Opcode::RegisterValue => {
1139 self.register_value(inst)?;
1140 continue;
1141 }
1142 // Where a frame is and what it returns to, built here for the same reason and one
1143 // more. The reason is the same: what the walk starts from is the frame pointer,
1144 // which is not a register a rule pattern can bind, and there is nothing in reading
1145 // the link the prologue saved that a proof over bitvectors could discharge. The
1146 // extra one is that how long the walk is comes out of a number beside the
1147 // instruction, so one of these is not one instruction but however many the depth
1148 // says, and a rule replaces a term with a term.
1149 Opcode::FrameAddress | Opcode::ReturnAddress => {
1150 self.frames(inst)?;
1151 continue;
1152 }
1153 // Built from the frame for the reason an `alloca` is, and from the convention for
1154 // the reason a call is: three of the four fields it writes are distances that do
1155 // not exist until the frame does, and the fourth is where the walk over the
1156 // argument registers stopped. A function that is not variadic has no such walk to
1157 // report, so it has nothing here and is refused below, which is the right answer
1158 // for a `va_start` in one.
1159 Opcode::VaStart if self.varargs.is_some() => {
1160 self.va_start(inst)?;
1161 continue;
1162 }
1163 // A return of more than one value, which is a structure small enough to come
1164 // back in a pair of registers. Built from the convention for the reason a call
1165 // is: which register each half goes in depends on the halves in front of it,
1166 // because the two register files are walked separately, and a pattern over a term
1167 // cannot see them. A return of one value is a term with a name and a rule, and it
1168 // stays one.
1169 //
1170 // A return of none in a function whose answer went through memory is here too,
1171 // and for a different reason: what it gives back is not written in the IR at all.
1172 // The convention says the address the caller handed over comes back, and only the
1173 // signature says this function was handed one.
1174 //
1175 // And a return of one eighty bit value, for a third reason: what a rule would
1176 // write is an instruction leaving the value in a register, and this one is left on
1177 // the x87 stack instead. A rule could not name that stack any more than any other
1178 // rule about this type could.
1179 Opcode::Return
1180 if self.source[self.source[inst].args].len() > 1
1181 || self.sret().is_some()
1182 || self.gives_back_x87(inst) =>
1183 {
1184 self.returned(inst)?;
1185 continue;
1186 }
1187 // A cast between a pointer and an integer of the same width, which on this
1188 // machine is every one the front end writes. No instruction at all, so no rule
1189 // could name one.
1190 Opcode::PtrToInt | Opcode::IntToPtr => {
1191 self.rename(inst)?;
1192 continue;
1193 }
1194 // A barrier, which is one instruction or none depending on the ordering. Written
1195 // by name because there is nothing about it a rule could be proved against, the
1196 // way there is nothing to prove about the address of a symbol.
1197 Opcode::Fence => {
1198 self.barrier(inst)?;
1199 continue;
1200 }
1201 // A hint, written by name for the reason a barrier is and one step further: not
1202 // only is there no equality for a proof to discharge, there is nothing about the
1203 // program around it either. Which of the four instructions it is comes out of the
1204 // number the builtin was given, which is beside the instruction rather than in it.
1205 Opcode::Prefetch => {
1206 self.hint(inst)?;
1207 continue;
1208 }
1209 // Stopping, written by name for the first half of the barrier's reason: it
1210 // computes nothing, so there is no term for a rule to replace, and what makes it
1211 // right is what the operating system does with the fault rather than anything a
1212 // proof over bitvectors could discharge.
1213 Opcode::Trap => {
1214 self.trap(inst);
1215 continue;
1216 }
1217 // A compare and exchange, which is written by name because it produces two values
1218 // and a rule produces one. The replacement of a rule is one term, a term names the
1219 // value an instruction computes, and there is no way in that language to say that
1220 // an instruction leaves an answer in one place and a yes or no in another.
1221 Opcode::Cmpxchg => {
1222 self.exchange(inst)?;
1223 continue;
1224 }
1225 // A read modify write, which is written by name for a different reason: it produces
1226 // one value, so a rule could name it, and what it does is not in the head a rule
1227 // matches on. Every one of the thirteen operations is the same opcode at the same
1228 // type and differs only in what is carried beside it, so one pattern would be all
1229 // thirteen patterns. Of the thirteen only the three with an instruction reach here,
1230 // since `crate::retry` turned the rest into loops a long way above this.
1231 Opcode::AtomicRmw => {
1232 self.modify(inst)?;
1233 continue;
1234 }
1235 // An `asm` statement, whose lowering is its template and there is no term for a
1236 // string. Written by name for the reason a barrier is, and before the x87 arm
1237 // below so that an `asm` holding a `long double` is refused as the `asm` it is
1238 // rather than as an instruction nothing computes.
1239 Opcode::InlineAsm => {
1240 self.assembly(inst)?;
1241 continue;
1242 }
1243 // Anything at all with an eighty bit float in it, which is the one arm here
1244 // chosen by a type rather than by an opcode, because what makes these different
1245 // is not what they do but where the value is. A `long double` has no register,
1246 // so it has no name in `crate::term` and no rule could bind one: every one of
1247 // these is a group of instructions over a frame slot, written out below.
1248 //
1249 // Last of the arms, so that a call and a return with one of these in them reach
1250 // the convention first and are refused by it, which is the truer answer: what is
1251 // wrong there is where the value has to travel and not that nothing can compute
1252 // it.
1253 _ if self.touches_x87(inst) => {
1254 self.x87(inst)?;
1255 continue;
1256 }
1257 _ => {}
1258 }
1259 let matched = matched.ok_or_else(|| self.unsupported(inst))?;
1260 self.emit(inst, &matched)?;
1261 // After it is built rather than when it matched, so that what is recorded is the rules
1262 // this function was lowered by and not the rules something was tried with.
1263 self.fired.mark(matched.rule);
1264 }
1265 // Whichever block the walk ended in rather than the one it started in. The two are the
1266 // same block for every function that does not save a place for a `__builtin_longjmp`, and
1267 // where they differ it is the last of them that the terminator and the arms belong to.
1268 // See [`Self::saves_place`].
1269 let last = self.at.expect("a block is being filled");
1270 self.edges(block, last)
1271 }
1272
1273 /// One call, which is built from the convention rather than matched against the table for the
1274 /// same reason the arguments of the function itself are.
1275 ///
1276 /// The arguments are read before the call is built, which is what materializes a constant
1277 /// argument into a register, since no call passes an immediate.
1278 ///
1279 /// A call to a name and a call through an address are both here, and what tells them apart is
1280 /// the opcode rather than whether a callee was recorded, which is the same thing the verifier
1281 /// reads. Through an address the first operand is the address and the arguments are the ones
1282 /// behind it, and everything after that is the same: where each argument goes, where the value
1283 /// comes back and which registers are gone across it are the convention's answers and the
1284 /// convention does not ask what is being called.
1285 fn called(&mut self, inst: Inst) -> Result<(), Unsupported> {
1286 let data = &self.source[inst];
1287 let Extra::Call(info) = data.extra else { return Err(self.unsupported(inst)) };
1288 let info = self.source[info];
1289 let indirect = data.opcode == Opcode::CallIndirect;
1290
1291 let values: Vec<Value> = self.source[data.args].to_vec();
1292 let callee = if indirect {
1293 let &address = values.first().ok_or_else(|| self.unsupported(inst))?;
1294 abi::Callee::Through(self.reg_of(address)?)
1295 } else {
1296 abi::Callee::Named(info.callee.ok_or_else(|| self.unsupported(inst))?)
1297 };
1298
1299 // What the ABI asks of each argument, read out before any of them is, because reading one
1300 // borrows the function this is a table in. The ones the signature names are the signature's
1301 // answer and the ones behind them are the call's, which is where a structure passed to a
1302 // variadic callee by value says that its bytes travel: there is no parameter to say it on.
1303 let signature = &self.source[info.signature];
1304 let variadic = signature.variadic;
1305 let named: Vec<Abi> = signature.params.iter().map(|param| param.abi).collect();
1306 let beyond: Vec<Abi> = self.source[info.varargs].to_vec();
1307 // Every value that comes back and not only the first. A structure small enough to travel
1308 // in registers comes back in up to two of them, and which register each half is in is the
1309 // convention's answer, which is why the whole list goes to the same place the arguments do
1310 // rather than to a rule.
1311 let returns: Vec<Type> = signature.return_types().collect();
1312
1313 let mut args = Vec::with_capacity(values.len());
1314 for (index, value) in values.into_iter().skip(usize::from(indirect)).enumerate() {
1315 let abi = named.get(index).or_else(|| beyond.get(index - named.len()));
1316 let abi = abi.copied().unwrap_or_default();
1317 let ty = self.source[value].ty;
1318 // What travels for an eighty bit value is its bytes, so what the call is handed is
1319 // where they are rather than a register they are in, and there is no register they
1320 // could be in. Everything else about it is a sixteen byte object passed by value and
1321 // is built by the same code.
1322 let reg =
1323 if abi::on_the_stack(ty) { self.x87_slot(value) } else { self.reg_of(value)? };
1324 args.push(abi::Passing { ty, reg, abi });
1325 }
1326 let block = self.at.expect("a block is being filled");
1327 let what = abi::Calling {
1328 callee,
1329 args: &args,
1330 returns: &returns,
1331 variadic,
1332 named: named.len(),
1333 at: self.source.span(inst),
1334 };
1335 let made = abi::call(&mut self.out, block, &what, self.conv, self.names)
1336 .map_err(|refused| Unsupported::Call { inst, refused })?;
1337 let calls = &mut self.stack.calls;
1338 *calls = Some(calls.unwrap_or(0).max(made.outgoing));
1339 // An eighty bit value came back on the x87 stack, and the one thing that has to happen
1340 // before anything else touches that stack is taking it off. So the `fstp` goes here, in
1341 // front of everything the block does next, and after it the value is in its slot and is
1342 // read the way every other one is.
1343 let results: Vec<Value> = self.source[inst].results().collect();
1344 if let [result] = results[..] {
1345 if abi::on_the_stack(self.source[result].ty) {
1346 let span = self.source.span(inst);
1347 let into = self.x87_slot(result);
1348 let into = self.through(into);
1349 self.x87_at("fstp_t", span, into);
1350 return Ok(());
1351 }
1352 }
1353 for (result, ®) in results.into_iter().zip(&made.results) {
1354 self.regs[result.index()] = Some(reg);
1355 }
1356 Ok(())
1357 }
1358
1359 /// The pointer a function returning through memory was handed, or nothing in a function that
1360 /// was not.
1361 ///
1362 /// It is the first parameter and the signature is what says so, since in the IR it is an
1363 /// ordinary pointer and reads like one everywhere in the body. A function with a signature
1364 /// like that and no entry block has nothing to give back and no body to give it back from.
1365 fn sret(&self) -> Option<Value> {
1366 let first = self.source.signature().params.first()?;
1367 if !matches!(first.abi, Abi::Sret { .. }) {
1368 return None;
1369 }
1370 self.source[self.source.entry()?].params.first().copied()
1371 }
1372
1373 /// One `return` the convention has to write, as the place each value has to be in by the end.
1374 ///
1375 /// One pseudo per value, each a read constrained to a return register, which is what a return
1376 /// of one value already is and is the whole of what either does. The `ret` itself comes from
1377 /// the epilogue for both, long after this, because the frame has to be given back first.
1378 ///
1379 /// The two register files are counted separately, so a structure of a `double` and a `long`
1380 /// leaves the `double` in the first vector register and the `long` in the first integer one
1381 /// rather than in the second of either. That is the same walk `rucc_codegen::abi` makes on
1382 /// the other side of the call, which is what makes the two ends agree.
1383 ///
1384 /// A function whose answer went through memory gives back the address it was handed, in front
1385 /// of nothing else, because a signature that returns that way returns nothing else. That the
1386 /// caller already knows the address is not enough: it is allowed to read the register instead,
1387 /// and a caller that does gets whatever the allocator last left there. In a leaf function that
1388 /// is usually the right answer by accident, and one call in the body is enough to make it a
1389 /// wild pointer, which is why this is written rather than left to luck.
1390 ///
1391 /// Where everything goes is worked out before anything is written, so a return this cannot
1392 /// make leaves no half of one behind.
1393 /// Whether what a `return` gives back is the one value that goes back on the x87 stack.
1394 fn gives_back_x87(&self, inst: Inst) -> bool {
1395 let [value] = self.source[self.source[inst].args] else { return false };
1396 abi::on_the_stack(self.source[value].ty)
1397 }
1398
1399 fn returned(&mut self, inst: Inst) -> Result<(), Unsupported> {
1400 let values: Vec<Value> = self.source[self.source[inst].args].to_vec();
1401 let (mut ints, mut floats) = (0usize, 0usize);
1402 let mut parts = Vec::with_capacity(values.len() + 1);
1403 // An eighty bit value goes back on the x87 stack, which is where the convention says it is
1404 // and is the one place a value is left rather than put in a register. So the whole of the
1405 // return is an `fld` of its slot, and the stack it leaves the value on is not empty at the
1406 // `ret`, which is the one time in this file that is true and is what the convention asks
1407 // for. What comes after is the epilogue, which gives the frame back and touches nothing in
1408 // the unit.
1409 if let [value] = values[..] {
1410 let ty = self.source[value].ty;
1411 if abi::on_the_stack(ty) && self.sret().is_none() {
1412 let span = self.source.span(inst);
1413 let from = self.x87_slot(value);
1414 let from = self.through(from);
1415 self.x87_at("fld_t", span, from);
1416 return Ok(());
1417 }
1418 }
1419 for value in self.sret().into_iter().chain(values) {
1420 let ty = self.source[value].ty;
1421 let at = if crate::term::in_vector_file(ty) { &mut floats } else { &mut ints };
1422 // Why it cannot come back, and not only that it cannot. A type that travels nowhere
1423 // says so itself, and a type that travels perfectly well ran out of registers.
1424 let missing = abi::refuses(ty).unwrap_or(Missing::NoRoom);
1425 let name = abi::ret_of(ty, *at).ok_or(Unsupported::Returned { inst, missing })?;
1426 *at += 1;
1427 // The register is the target's answer and not one worked out here, the same as it is
1428 // for a return of one value, so that both halves of a pair and every rule that writes
1429 // half of one are reading the same table.
1430 let opcode = name.strip_prefix(PREFIX).expect("a machine instruction of this target");
1431 let form = x86_64::form(opcode).ok_or_else(|| self.unsupported(inst))?;
1432 let [desc] = form.operands() else { return Err(self.unsupported(inst)) };
1433 parts.push((self.names.intern(name), self.reg_of(value)?, *desc));
1434 }
1435
1436 let block = self.at.expect("a block is being filled");
1437 let span = self.source.span(inst);
1438 for (opcode, reg, desc) in parts {
1439 let operand = mir::Operand {
1440 reg,
1441 class: desc.class,
1442 role: desc.role,
1443 constraint: desc.constraint,
1444 };
1445 self.out.build(block, mir::Opcode::new(opcode)).at(span).operand(operand).finish();
1446 }
1447 Ok(())
1448 }
1449
1450 /// One `alloca`: the bytes it asks for go on the list the frame is laid out from, and the
1451 /// address of them is one instruction.
1452 ///
1453 /// The instruction is a `lea` off the stack pointer, which is the one register that reaches
1454 /// the frame in every function, and its displacement is left at nothing because there is no
1455 /// frame yet. Which instruction is waiting for which local is remembered, and
1456 /// [`crate::finish`] fills the numbers in after [`crate::frame::Frame`] has placed them.
1457 ///
1458 /// There is deliberately no rule for `alloca` and no name for one in [`crate::term`], and
1459 /// that is what stops it being folded into something else. An operand shown as the
1460 /// instruction that computed it is offered to the matcher by its name, so an `alloca` with no
1461 /// name is one no pattern can reach past, and the address it computes is always in a register
1462 /// by the time anything reads it.
1463 fn reserve(&mut self, inst: Inst) -> Result<(), Unsupported> {
1464 let data = &self.source[inst];
1465 // A variable length array carries the size it wants as an operand rather than in the
1466 // instruction, which is the whole of what tells the two apart here.
1467 if let Some(&size) = self.source[data.args].first() {
1468 return self.grow(inst, size);
1469 }
1470 let Extra::Mem(mem) = data.extra else { return Err(self.unsupported(inst)) };
1471 let info = self.source[mem];
1472 let size = u32::try_from(info.size)
1473 .map_err(|_| Unsupported::Dynamic { inst, growing: Growing::Huge })?;
1474 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
1475
1476 // At least one, because the frame divides by the alignment and an object with no
1477 // alignment at all is one the front end had nothing to say about rather than one that may
1478 // go anywhere.
1479 let index = self.stack.locals.len();
1480 self.stack.locals.push(Local { size, align: info.align.max(1) });
1481 if let Some(decl) = self.source.mem_decl(mem) {
1482 self.stack.declared.push((index, decl));
1483 }
1484
1485 let block = self.at.expect("a block is being filled");
1486 let reg = self.new_reg(result);
1487 let span = self.source.span(inst);
1488 let lea = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.lea)));
1489 let sp = mir::Operand::read(mir::Reg::physical(self.conv.stack_pointer), self.gpr);
1490 let made =
1491 self.out.build(block, lea).at(span).def(reg, self.gpr).mem(mir::Mem::at(sp)).finish();
1492 self.stack.addresses.push((made, index));
1493 Ok(())
1494 }
1495
1496 /// The other kind of `alloca`: one whose size the function does not know until it runs, which
1497 /// is what a variable length array is.
1498 ///
1499 /// Nothing about it is a slot the frame laid out, because the frame is laid out once and this
1500 /// happens as often as control reaches the declaration. The bytes come off the stack pointer
1501 /// where the declaration stands, which is two instructions:
1502 ///
1503 /// ```text
1504 /// sub sp, bytes the stack pointer moves down over the memory, which is what takes it
1505 /// lea reg, [sp+n] where the memory starts, which is above the outgoing argument area
1506 /// ```
1507 ///
1508 /// The displacement is left at nothing for the reason the constant kind leaves its own at
1509 /// nothing, and for a different number: that area belongs to the arguments of whatever this
1510 /// function calls, it stays at the bottom of the frame wherever the bottom has moved to, and
1511 /// how big it is is not known until every call in the function has been seen.
1512 ///
1513 /// The bytes are already a multiple of the stack pointer's alignment by the time they arrive,
1514 /// because [`crate::expand::rounds`] rounded them up in the IR, so nothing here has to mask the
1515 /// stack pointer afterwards and the stack pointer stays somewhere a call can be made from.
1516 ///
1517 /// Two instructions here and not always two in the finished function. On a command line that
1518 /// asked for the stack to be touched a page at a time, the subtraction becomes a loop that
1519 /// walks the same distance a page at a time, which [`crate::finish`] writes. That is why the
1520 /// instruction is written down in [`Stack::grown`] as well as left where it is.
1521 ///
1522 /// An array wanting more alignment than the convention leaves the stack pointer with does not
1523 /// reach here asking for it: [`crate::expand::rounds`] gives it the alignment in extra bytes
1524 /// and turns the array into a `ptr_add` of the offset that lands inside them, so what arrives
1525 /// is a block asking for the convention's alignment like any other. The refusal below is what
1526 /// answers IR that came from somewhere other than that pass, since forcing the alignment here
1527 /// would be a second rounding of a register the frame already rounded, and after it no
1528 /// constant reaches the rest of the frame from anywhere. See `Growing` in [`crate::frame`].
1529 fn grow(&mut self, inst: Inst, size: Value) -> Result<(), Unsupported> {
1530 let data = &self.source[inst];
1531 let Extra::Mem(mem) = data.extra else { return Err(self.unsupported(inst)) };
1532 let info = self.source[mem];
1533 if info.align > self.conv.stack_align {
1534 return Err(Unsupported::Dynamic { inst, growing: Growing::Aligned });
1535 }
1536 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
1537 let bytes = self.reg_of(size)?;
1538
1539 let block = self.at.expect("a block is being filled");
1540 let span = self.source.span(inst);
1541 let stack = mir::Reg::physical(self.conv.stack_pointer);
1542 let grow = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.grow)));
1543 let took = self
1544 .out
1545 .build(block, grow)
1546 .at(span)
1547 .operand(mir::Operand::write(stack, self.gpr))
1548 .operand(mir::Operand::read(stack, self.gpr))
1549 .operand(mir::Operand::read(bytes, self.gpr))
1550 .finish();
1551 self.stack.grown.push(took);
1552
1553 let reg = self.new_reg(result);
1554 let lea = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.lea)));
1555 let sp = mir::Operand::read(stack, self.gpr);
1556 let made =
1557 self.out.build(block, lea).at(span).def(reg, self.gpr).mem(mir::Mem::at(sp)).finish();
1558 self.stack.dynamic.push(made);
1559 self.stack.grown_at.get_or_insert(inst);
1560 Ok(())
1561 }
1562
1563 /// Where the stack pointer is, kept so that something later can put it back.
1564 ///
1565 /// One move out of the stack pointer and one move into it, which is the whole of what the two
1566 /// halves are. What makes them worth writing is where the front end puts them: a scope holding
1567 /// a variable length array saves the stack pointer as it opens and puts it back as it closes,
1568 /// so a loop declaring one takes its bytes once round rather than once per iteration, and a
1569 /// jump out of the scope gives the bytes back on the way out.
1570 ///
1571 /// The value travels in an ordinary register the allocator hands out, so it may be spilled like
1572 /// any other, and a spill slot in a frame that grows is reached through the frame pointer,
1573 /// which is exactly the register that still means something after the stack pointer has moved.
1574 fn stack_pointer(&mut self, inst: Inst, into: bool) -> Result<(), Unsupported> {
1575 let data = &self.source[inst];
1576 let block = self.at.expect("a block is being filled");
1577 let span = self.source.span(inst);
1578 let stack = mir::Reg::physical(self.conv.stack_pointer);
1579 let mov = x86_64::FRAME.moves(self.gpr).expect("a class the target says how to move").mov;
1580 let mov = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{mov}")));
1581 let (write, read) = if into {
1582 let &saved = self.source[data.args].first().ok_or_else(|| self.unsupported(inst))?;
1583 (stack, self.reg_of(saved)?)
1584 } else {
1585 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
1586 (self.new_reg(result), stack)
1587 };
1588 self.out
1589 .build(block, mov)
1590 .at(span)
1591 .operand(mir::Operand::write(write, self.gpr))
1592 .operand(mir::Operand::read(read, self.gpr))
1593 .finish();
1594 // Only the write is a move of the stack pointer, and it is the one that makes the frame a
1595 // growing one. A read of it in a function that never writes it back is a function that
1596 // asked where the stack was and did nothing with the answer.
1597 if into {
1598 self.stack.grown_at.get_or_insert(inst);
1599 }
1600 Ok(())
1601 }
1602
1603 /// Whether an instruction has an eighty bit float anywhere in it.
1604 ///
1605 /// Producing one and reading one are the same question here, because what makes one of these
1606 /// different from every other instruction is not the operation but where the value is. A
1607 /// `long double` is on the x87 stack while it is being worked on and in a frame slot the rest
1608 /// of the time, and neither of those is somewhere the operand of a rule could point.
1609 fn touches_x87(&self, inst: Inst) -> bool {
1610 let data = &self.source[inst];
1611 data.results().any(|value| on_x87(self.source[value].ty))
1612 || self.source[data.args].iter().any(|&arg| on_x87(self.source[arg].ty))
1613 }
1614
1615 /// Everything that happens to an eighty bit float, as the group of instructions it is.
1616 ///
1617 /// The first six move one, and every one of those is a load, a store, or a load and a store at
1618 /// two different formats, because that is the whole of what this machine converts with: the
1619 /// x87 has no instruction that turns one thing on its stack into another, so a widening is
1620 /// `fld` of the narrow format and a narrowing is `fstp` of it.
1621 ///
1622 /// The rest work on one, and they are here rather than in a rule for the same reason the six
1623 /// are. An add is a push, a push, the add and a pop, and what passes between those four is the
1624 /// top of a stack nothing allocates from, so there is no value in the middle of the group for
1625 /// a pattern to bind or a replacement to name. The comparison is the same shape with its last
1626 /// two instructions folded into one opcode, which is where the byte it produces comes from.
1627 ///
1628 /// Every group leaves the stack as empty as it found it, which is what `spec/10-backend.md`
1629 /// section 10.8 asks of one and is why nothing in this file has to track a depth: each push
1630 /// below is answered by a pop a line or two later, so no two groups can ever be looking at
1631 /// the same eight registers.
1632 fn x87(&mut self, inst: Inst) -> Result<(), Unsupported> {
1633 match self.source[inst].opcode {
1634 Opcode::Load => self.x87_load(inst),
1635 Opcode::Store => self.x87_store(inst),
1636 Opcode::FPExt => self.x87_widen(inst),
1637 Opcode::FPTrunc => self.x87_narrow(inst),
1638 Opcode::SIToFP => self.x87_from_signed(inst),
1639 Opcode::FPToSI => self.x87_to_signed(inst),
1640 Opcode::FAdd => self.x87_arith(inst, "fadd_p"),
1641 Opcode::FSub => self.x87_arith(inst, "fsubr_p"),
1642 Opcode::FMul => self.x87_arith(inst, "fmul_p"),
1643 Opcode::FDiv => self.x87_arith(inst, "fdivr_p"),
1644 Opcode::FNeg => self.x87_flip(inst),
1645 Opcode::FCmp => self.x87_compare(inst),
1646 Opcode::FConst => self.x87_const(inst),
1647 _ => Err(self.unsupported(inst)),
1648 }
1649 }
1650
1651 /// The eighty bit parameters of a block, copied out of the addresses an edge handed over and
1652 /// into slots of the block's own.
1653 ///
1654 /// What crosses an edge for a value of this type is an address, because the value is sixteen
1655 /// bytes of the frame and no register holds any of it. The block cannot keep that address: a
1656 /// second edge into the same block hands over a second one, and a read after the block would
1657 /// then be a read of whichever edge was taken rather than of one place. So the block has a
1658 /// slot per parameter and the bytes are copied into it here, which is the move on an edge that
1659 /// every other type gets from the allocator.
1660 ///
1661 /// Every load runs before every store and the stores run backwards, so all of the values are
1662 /// on the x87 stack at once and nothing reads a slot another one has already written. That
1663 /// costs nothing in the ordinary case of one parameter and is what makes the back edge of a
1664 /// loop that swaps two of these work. It is also the reason for the limit: the stack is eight
1665 /// deep, and a block with more of these than that is refused rather than copied in an order
1666 /// that could be wrong.
1667 fn settle(&mut self, block: Block, arriving: &[(Value, mir::Reg)]) -> Result<(), Unsupported> {
1668 let Some(&(first, _)) = arriving.first() else { return Ok(()) };
1669 if arriving.len() > X87_DEPTH {
1670 let ty = self.source[first].ty;
1671 return Err(Unsupported::Phi { block, count: arriving.len(), ty });
1672 }
1673 // A block parameter comes from no instruction, so what this points at is the first thing
1674 // in the block, which is where a reader looking for the copy would look.
1675 let first_inst = self.source.insts(block).next();
1676 let span = first_inst.map_or(Span::DUMMY, |it| self.source.span(it));
1677 for &(_, reg) in arriving {
1678 let from = self.through(reg);
1679 self.x87_at("fld_t", span, from);
1680 }
1681 for &(param, _) in arriving.iter().rev() {
1682 let into = self.x87_slot(param);
1683 let into = self.through(into);
1684 self.x87_at("fstp_t", span, into);
1685 }
1686 Ok(())
1687 }
1688
1689 /// The frame slot an eighty bit value lives in, as its address in a fresh register.
1690 ///
1691 /// The slot is the value's for the whole function and is taken the first time somebody asks.
1692 /// The address is worked out again every time, which is a `lea` per use and is deliberate: one
1693 /// address kept in a register from the definition to the last use would hold a general purpose
1694 /// register open across everything in between, and a function with a handful of these in it
1695 /// would spend its registers on addresses of things rather than on things.
1696 fn x87_slot(&mut self, value: Value) -> mir::Reg {
1697 // An argument of the function has a slot already and it is the caller's. The convention
1698 // puts the bytes in the argument area and hands over where they are, so the address that
1699 // arrived is the answer and no second copy of the value is made. Nothing ever writes to a
1700 // value of this type once it exists, so nothing writes to the caller's copy either. A
1701 // parameter of any other block is not this: what arrived there is an address a predecessor
1702 // chose, [`Lowering::settle`] has already copied the bytes out of it, and the slot those
1703 // bytes landed in is the one below.
1704 let entry = self.source.entry();
1705 if let (Def::Param { block, .. }, Some(reg)) =
1706 (self.source[value].def, self.regs[value.index()])
1707 {
1708 if entry == Some(block) {
1709 return reg;
1710 }
1711 }
1712 let index = match self.slots[value.index()] {
1713 Some(index) => index,
1714 None => {
1715 let index = self.stack.locals.len();
1716 self.stack.locals.push(Local { size: X87_BYTES, align: X87_BYTES });
1717 self.slots[value.index()] = Some(index);
1718 index
1719 }
1720 };
1721 let block = self.at.expect("a block is being filled");
1722 self.frame_address(block, index)
1723 }
1724
1725 /// The bytes a value crosses between a register and the x87 stack through, as their address
1726 /// in a fresh register.
1727 fn x87_crossing(&mut self) -> mir::Reg {
1728 let index = match self.crossing {
1729 Some(index) => index,
1730 None => {
1731 let index = self.stack.locals.len();
1732 self.stack.locals.push(Local { size: X87_CROSSING, align: X87_CROSSING });
1733 self.crossing = Some(index);
1734 index
1735 }
1736 };
1737 let block = self.at.expect("a block is being filled");
1738 self.frame_address(block, index)
1739 }
1740
1741 /// The two control words, as the address of the first of them in a fresh register.
1742 fn x87_control(&mut self) -> mir::Reg {
1743 let index = match self.control {
1744 Some(index) => index,
1745 None => {
1746 let index = self.stack.locals.len();
1747 self.stack.locals.push(Local { size: 4, align: 4 });
1748 self.control = Some(index);
1749 index
1750 }
1751 };
1752 let block = self.at.expect("a block is being filled");
1753 self.frame_address(block, index)
1754 }
1755
1756 /// An address held in a register, as the addressing mode that reaches it.
1757 fn through(&self, reg: mir::Reg) -> mir::Mem {
1758 mir::Mem::at(mir::Operand::read(reg, self.gpr))
1759 }
1760
1761 /// One instruction of a group, which names an address and nothing else.
1762 ///
1763 /// Every x87 instruction that moves a value is one of these. What it does to the stack is in
1764 /// the mnemonic rather than in an operand, so there is no register to write down and no
1765 /// register the allocator gets a say in.
1766 fn x87_at(&mut self, name: &str, span: Span, at: mir::Mem) {
1767 let block = self.at.expect("a block is being filled");
1768 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
1769 self.out.build(block, opcode).at(span).mem(at).finish();
1770 }
1771
1772 /// The one instruction of a group that reaches the program's own memory.
1773 ///
1774 /// A `long double` moves in two instructions with a frame slot at one end of them, and the
1775 /// other end is the address the program wrote. That end is the access, so it is the one that
1776 /// carries what the program said about it, and the trip through the slot is this compiler's
1777 /// own business the way a spill is. See [`Self::carried`].
1778 fn x87_touching(&mut self, name: &str, inst: Inst, at: mir::Mem) {
1779 let block = self.at.expect("a block is being filled");
1780 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
1781 let (span, flags) = (self.source.span(inst), self.carried(inst));
1782 self.out.build(block, opcode).at(span).flags(flags).mem(at).finish();
1783 }
1784
1785 /// One instruction of a group that names nothing at all.
1786 ///
1787 /// The arithmetic is these. Both of an add's operands are already on the stack when it runs
1788 /// and so is where the answer goes, and the stack is not somewhere an instruction says, so
1789 /// `faddp` has an argument in the assembler's syntax and nothing here for the argument to come
1790 /// from. What it works on is which two pushes came before it, which is a fact about the order
1791 /// of the group and is why the group is written in one place.
1792 fn x87_only(&mut self, name: &str, span: Span) {
1793 let block = self.at.expect("a block is being filled");
1794 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
1795 self.out.build(block, opcode).at(span).finish();
1796 }
1797
1798 /// A `load` of a `long double`: onto the stack from where it was, and off it into the slot.
1799 ///
1800 /// Two instructions rather than the two general purpose moves the same sixteen bytes would
1801 /// take, because `fld` and `fstp` at this format neither convert nor look: the value goes on
1802 /// in the format it was already in and comes back off in it, so a signalling NaN stays one
1803 /// and nothing is raised. Which is what makes this a copy at all.
1804 fn x87_load(&mut self, inst: Inst) -> Result<(), Unsupported> {
1805 let (args, result) = self.ends(inst)?;
1806 let &address = args.first().ok_or_else(|| self.unsupported(inst))?;
1807 let span = self.source.span(inst);
1808 let from = self.reg_of(address)?;
1809 let from = self.through(from);
1810 let into = self.x87_slot(result);
1811 let into = self.through(into);
1812 self.x87_touching("fld_t", inst, from);
1813 self.x87_at("fstp_t", span, into);
1814 Ok(())
1815 }
1816
1817 /// A `store` of a `long double`: the same pair the other way round.
1818 fn x87_store(&mut self, inst: Inst) -> Result<(), Unsupported> {
1819 let args = self.source[self.source[inst].args].to_vec();
1820 let [value, address] = args[..] else { return Err(self.unsupported(inst)) };
1821 let span = self.source.span(inst);
1822 let from = self.x87_slot(value);
1823 let from = self.through(from);
1824 let into = self.reg_of(address)?;
1825 let into = self.through(into);
1826 self.x87_at("fld_t", span, from);
1827 self.x87_touching("fstp_t", inst, into);
1828 Ok(())
1829 }
1830
1831 /// A `float`, a `double` or an integer becoming a `long double`.
1832 ///
1833 /// Through memory, because the x87 reads memory and nothing else: the value is in a register
1834 /// the machine has and the unit has no way to be handed one, so it is written to the crossing
1835 /// bytes and loaded back at the format that widens it. Every one of these is exact. Sixty four
1836 /// bits of significand and fifteen of exponent hold every `float`, every `double` and every
1837 /// sixty four bit integer outright, so none of the four can round and none can raise.
1838 fn x87_across(
1839 &mut self,
1840 inst: Inst,
1841 put: &'static str,
1842 class: RegClass,
1843 get: &'static str,
1844 ) -> Result<(), Unsupported> {
1845 let (args, result) = self.ends(inst)?;
1846 let &source = args.first().ok_or_else(|| self.unsupported(inst))?;
1847 let span = self.source.span(inst);
1848 let value = self.reg_of(source)?;
1849 let across = self.x87_crossing();
1850 let across = self.through(across);
1851 let into = self.x87_slot(result);
1852 let into = self.through(into);
1853
1854 let block = self.at.expect("a block is being filled");
1855 let store = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{put}")));
1856 self.out.build(block, store).at(span).uses(value, class).mem(across).finish();
1857 self.x87_at(get, span, across);
1858 self.x87_at("fstp_t", span, into);
1859 Ok(())
1860 }
1861
1862 /// A `long double` becoming a `float`, a `double` or an integer.
1863 ///
1864 /// Through memory for the reason above and in the same three instructions backwards. The two
1865 /// that go to a float round to nearest, which is what the control word says unless somebody
1866 /// has changed it and is what C wants. The two that go to an integer do not, which is why they
1867 /// do not come here.
1868 fn x87_back(
1869 &mut self,
1870 inst: Inst,
1871 put: &'static str,
1872 get: &'static str,
1873 class: RegClass,
1874 ) -> Result<(), Unsupported> {
1875 let (args, result) = self.ends(inst)?;
1876 let &source = args.first().ok_or_else(|| self.unsupported(inst))?;
1877 let span = self.source.span(inst);
1878 let from = self.x87_slot(source);
1879 let from = self.through(from);
1880 let across = self.x87_crossing();
1881 let across = self.through(across);
1882
1883 self.x87_at("fld_t", span, from);
1884 self.x87_at(put, span, across);
1885 let block = self.at.expect("a block is being filled");
1886 let reg = self.new_reg(result);
1887 let load = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{get}")));
1888 self.out.build(block, load).at(span).def(reg, class).mem(across).finish();
1889 Ok(())
1890 }
1891
1892 /// An `fpext` up to a `long double`, which is the only direction this machine has one in.
1893 fn x87_widen(&mut self, inst: Inst) -> Result<(), Unsupported> {
1894 let sse = self.conv.sse_class;
1895 match self.source[self.narrow(inst)?].ty.bits() {
1896 32 => self.x87_across(inst, "movss_mr", sse, "fld_s"),
1897 64 => self.x87_across(inst, "movsd_mr", sse, "fld_l"),
1898 _ => Err(self.unsupported(inst)),
1899 }
1900 }
1901
1902 /// An `fptrunc` down from a `long double`, which is the other direction of the same.
1903 fn x87_narrow(&mut self, inst: Inst) -> Result<(), Unsupported> {
1904 let sse = self.conv.sse_class;
1905 let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
1906 match self.source[result].ty.bits() {
1907 32 => self.x87_back(inst, "fstp_s", "movss_rm", sse),
1908 64 => self.x87_back(inst, "fstp_l", "movsd_rm", sse),
1909 _ => Err(self.unsupported(inst)),
1910 }
1911 }
1912
1913 /// A `sitofp` up to a `long double`.
1914 ///
1915 /// Thirty two bits and sixty four, and nothing narrower, because C widens an integer to `int`
1916 /// before it converts one and the front end writes that widening down. An unsigned integer is
1917 /// not here at all: `fild` reads its operand as signed, so a value above the signed range
1918 /// comes back short by two to the sixty fourth and has to be added back, which is arithmetic
1919 /// rather than a move and waits with the rest of it.
1920 fn x87_from_signed(&mut self, inst: Inst) -> Result<(), Unsupported> {
1921 let gpr = self.gpr;
1922 match self.source[self.narrow(inst)?].ty.bits() {
1923 32 => self.x87_across(inst, "mov_mr_32", gpr, "fild_l"),
1924 64 => self.x87_across(inst, "mov_mr_64", gpr, "fild_ll"),
1925 _ => Err(self.unsupported(inst)),
1926 }
1927 }
1928
1929 /// An `fptosi` down from a `long double`, which is the one conversion here with no single
1930 /// instruction behind it.
1931 ///
1932 /// C cuts towards zero and the unit rounds the way its control word says, so the store that
1933 /// takes the value off the stack is wrapped in the control word being saved, changed and put
1934 /// back. Five instructions around the one that does the work, and three more moving the word
1935 /// through a register, because this machine has no way to OR a constant into memory at this
1936 /// width. The unit has a shorter answer in `fisttp`, and `spec/10-backend.md` section 10.8
1937 /// says why it is not used: it is SSE3, the x86-64 baseline is not, and there is nothing here
1938 /// that can gate an instruction on a feature yet.
1939 fn x87_to_signed(&mut self, inst: Inst) -> Result<(), Unsupported> {
1940 let (args, result) = self.ends(inst)?;
1941 let &source = args.first().ok_or_else(|| self.unsupported(inst))?;
1942 let (put, get) = match self.source[result].ty.bits() {
1943 32 => ("fistp_l", "mov_rm_32"),
1944 64 => ("fistp_ll", "mov_rm_64"),
1945 _ => return Err(self.unsupported(inst)),
1946 };
1947 let span = self.source.span(inst);
1948 let gpr = self.gpr;
1949 let from = self.x87_slot(source);
1950 let from = self.through(from);
1951 let across = self.x87_crossing();
1952 let across = self.through(across);
1953 let control = self.x87_control();
1954 let saved = self.through(control).plus(0);
1955 let cut = self.through(control).plus(2);
1956
1957 // The word the unit has now, into the first of the two slots and into a register, with the
1958 // rounding field turned to truncate on the way to the second.
1959 self.x87_at("fnstcw", span, saved);
1960 let block = self.at.expect("a block is being filled");
1961 let was = self.out.new_vreg(gpr);
1962 let read = mir::Opcode::new(self.names.intern("x64.mov_rm_16"));
1963 self.out.build(block, read).at(span).def(was, gpr).mem(saved).finish();
1964 let now = self.out.new_vreg(gpr);
1965 let set = mir::Opcode::new(self.names.intern("x64.or_ri_16"));
1966 // Two address, which is written out here rather than taken from the two shorthands
1967 // because the shorthands leave an operand unconstrained: this machine ORs into the
1968 // register it read, so the two have to be the same one and only the constraint says so.
1969 self.out
1970 .build(block, set)
1971 .at(span)
1972 .operand(mir::Operand::write(now, gpr).with(Constraint::Reuse(1)))
1973 .operand(mir::Operand::read(was, gpr))
1974 .imm(X87_TRUNCATE)
1975 .finish();
1976 let write = mir::Opcode::new(self.names.intern("x64.mov_mr_16"));
1977 self.out.build(block, write).at(span).uses(now, gpr).mem(cut).finish();
1978
1979 // The conversion itself, under the changed word, and then the word the unit had put back
1980 // before anything else runs.
1981 self.x87_at("fldcw", span, cut);
1982 self.x87_at("fld_t", span, from);
1983 self.x87_at(put, span, across);
1984 self.x87_at("fldcw", span, saved);
1985
1986 let block = self.at.expect("a block is being filled");
1987 let reg = self.new_reg(result);
1988 let load = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{get}")));
1989 self.out.build(block, load).at(span).def(reg, gpr).mem(across).finish();
1990 Ok(())
1991 }
1992
1993 /// A constant of this type, as the bits of it written into its slot.
1994 ///
1995 /// No x87 instruction at all, which is the surprise here. A slot holding an eighty bit value is
1996 /// the value, so a constant is ten bytes put where the value lives, and the unit never has to
1997 /// see it: whatever reads it will `fld` it out of the slot the way it reads any other one.
1998 ///
1999 /// Ten bytes in two goes, because the machine stores eight at a time and there is no store of
2000 /// an immediate to memory, so each half is put in a register first. The six bytes above the ten
2001 /// are left alone, since nothing reads them: they are the padding that makes the type sixteen
2002 /// wide and they are unspecified in the psABI rather than zero.
2003 ///
2004 /// The other way is a constant pool, an `fldt` of a symbol, and a relocation, which is what a
2005 /// compiler with somewhere to put a literal does. This back end has nowhere to put one yet, and
2006 /// four instructions in the frame is what that costs until it does.
2007 fn x87_const(&mut self, inst: Inst) -> Result<(), Unsupported> {
2008 let Extra::Imm(imm) = self.source[inst].extra else { return Err(self.unsupported(inst)) };
2009 let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
2010 let bits = self.source[imm].bits();
2011 let span = self.source.span(inst);
2012 let gpr = self.gpr;
2013 let slot = self.x87_slot(result);
2014 let low = self.through(slot).plus(0);
2015 let high = self.through(slot).plus(8);
2016
2017 let block = self.at.expect("a block is being filled");
2018 for (bytes, at, into) in
2019 [(bits as u64 as i64, low, "64"), (((bits >> 64) & 0xffff) as i64, high, "16")]
2020 {
2021 let held = self.out.new_vreg(gpr);
2022 let put = mir::Opcode::new(self.names.intern(&format!("{PREFIX}mov_ri_{into}")));
2023 self.out.build(block, put).at(span).def(held, gpr).imm(bytes).finish();
2024 let store = mir::Opcode::new(self.names.intern(&format!("{PREFIX}mov_mr_{into}")));
2025 self.out.build(block, store).at(span).uses(held, gpr).mem(at).finish();
2026 }
2027 Ok(())
2028 }
2029
2030 /// One arithmetic instruction on two eighty bit values, as the four it takes.
2031 ///
2032 /// The left operand is pushed first and the right one on top of it, so the left ends up
2033 /// underneath and the answer wanted is the one below against the top in that order. Which of
2034 /// the two mnemonics computes that is a question about the spelling rather than about the
2035 /// machine, and the two spellings disagree. Intel's `FSUBP ST(i), ST(0)` is `ST(i) - ST(0)`
2036 /// and is `DE E8+i`, and AT&T's `fsubp` is `DE E0+i`, which is the other subtraction. This
2037 /// compiler writes AT&T and encodes what gas encodes, so what it asks for here is `fsubr_p`
2038 /// and `fdivr_p`, and the `r` is not a reversal of anything the code generator decided.
2039 ///
2040 /// An addition and a multiplication have one form each and do not care, which is why a test
2041 /// that reads the mnemonic back would not have caught this and one that computes a subtraction
2042 /// and checks the answer does.
2043 ///
2044 /// The answer is left where the deeper of the two was and the shallower is gone, which is what
2045 /// the `p` on the mnemonic means, so one push has already been paid back by the time the
2046 /// `fstp` runs and the stack is level again after it.
2047 ///
2048 /// Nothing here is folded and nothing is reused. Two values that are the same value get two
2049 /// pushes of the same slot, and an operand that was just computed is read back out of the slot
2050 /// it was written to rather than left on the stack, which costs a store and a load per
2051 /// instruction in an expression. Keeping a partial result on the stack across the next
2052 /// instruction's operands means knowing how deep the stack is at every point in the block, and
2053 /// that is a different thing from writing a group.
2054 fn x87_arith(&mut self, inst: Inst, with: &'static str) -> Result<(), Unsupported> {
2055 let (args, result) = self.ends(inst)?;
2056 let [left, right] = args[..] else { return Err(self.unsupported(inst)) };
2057 let span = self.source.span(inst);
2058 let left = self.x87_slot(left);
2059 let left = self.through(left);
2060 let right = self.x87_slot(right);
2061 let right = self.through(right);
2062 let into = self.x87_slot(result);
2063 let into = self.through(into);
2064 self.x87_at("fld_t", span, left);
2065 self.x87_at("fld_t", span, right);
2066 self.x87_only(with, span);
2067 self.x87_at("fstp_t", span, into);
2068 Ok(())
2069 }
2070
2071 /// A negation, which is a push, the sign bit turned over and a pop.
2072 ///
2073 /// `fchs` does not read the value as a number, so this is right for a zero, for an infinity
2074 /// and for a NaN, and it raises nothing on any of them. Which is what C asks of a negation and
2075 /// is not what subtracting from zero would give: `0.0L - x` is a different answer at a
2076 /// negative zero and a signalling one at a NaN.
2077 fn x87_flip(&mut self, inst: Inst) -> Result<(), Unsupported> {
2078 let (args, result) = self.ends(inst)?;
2079 let &source = args.first().ok_or_else(|| self.unsupported(inst))?;
2080 let span = self.source.span(inst);
2081 let from = self.x87_slot(source);
2082 let from = self.through(from);
2083 let into = self.x87_slot(result);
2084 let into = self.through(into);
2085 self.x87_at("fld_t", span, from);
2086 self.x87_only("fchs", span);
2087 self.x87_at("fstp_t", span, into);
2088 Ok(())
2089 }
2090
2091 /// A comparison of two eighty bit values, as the two pushes and the one opcode that reads them.
2092 ///
2093 /// The right operand is pushed first and the left one on top of it, which is the other way
2094 /// round from the arithmetic and is because `fucomip` asks about the top against what is under
2095 /// it: the comparison this machine can do is the top's, so the value the predicate is about
2096 /// has to be the top. The pop that gets the loser off the stack and the byte that reads the
2097 /// flags are both inside the opcode, since what passes between those and the comparison is the
2098 /// flags and the flags are not something anything here can name.
2099 ///
2100 /// Which of the ten opcodes, and which way round, is the same table the vector comparisons
2101 /// match against in `rules/x86-64.rules`, and it has to stay the same table: a predicate that
2102 /// picked a different condition here than there would be a `long double` comparison that
2103 /// disagreed with the `double` comparison of the same two numbers, which is the one thing a
2104 /// wider format is not allowed to do.
2105 ///
2106 /// The always false and the always true are refused rather than folded into a constant,
2107 /// because a comparison this machine never has to do is one the optimizer should have removed
2108 /// and an instruction here that quietly agreed with it would hide that it did not.
2109 fn x87_compare(&mut self, inst: Inst) -> Result<(), Unsupported> {
2110 let Extra::FloatPred(pred) = self.source[inst].extra else {
2111 return Err(self.unsupported(inst));
2112 };
2113 let (args, result) = self.ends(inst)?;
2114 let [left, right] = args[..] else { return Err(self.unsupported(inst)) };
2115 // Two of the fourteen need a second byte and an instruction to put the two together,
2116 // because they are two conditions at once: an ordered equal is equal and not unordered,
2117 // and an unordered not equal is either. The opcode carries all of that and says here only
2118 // that it writes somewhere else as well.
2119 let (name, reversed, both) = match pred {
2120 FloatPred::Ogt => ("fucomip_set_a", false, false),
2121 FloatPred::Oge => ("fucomip_set_ae", false, false),
2122 FloatPred::Olt => ("fucomip_set_a", true, false),
2123 FloatPred::Ole => ("fucomip_set_ae", true, false),
2124 FloatPred::One => ("fucomip_set_ne", false, false),
2125 FloatPred::Ord => ("fucomip_set_np", false, false),
2126 FloatPred::Uno => ("fucomip_set_p", false, false),
2127 FloatPred::Ueq => ("fucomip_set_e", false, false),
2128 FloatPred::Ult => ("fucomip_set_b", false, false),
2129 FloatPred::Ule => ("fucomip_set_be", false, false),
2130 FloatPred::Ugt => ("fucomip_set_b", true, false),
2131 FloatPred::Uge => ("fucomip_set_be", true, false),
2132 FloatPred::Oeq => ("fucomip_set_e_and_np", false, true),
2133 FloatPred::Une => ("fucomip_set_ne_or_p", false, true),
2134 FloatPred::False | FloatPred::True => return Err(self.unsupported(inst)),
2135 };
2136 let (top, under) = if reversed { (right, left) } else { (left, right) };
2137
2138 let span = self.source.span(inst);
2139 let gpr = self.gpr;
2140 let under = self.x87_slot(under);
2141 let under = self.through(under);
2142 let top = self.x87_slot(top);
2143 let top = self.through(top);
2144 self.x87_at("fld_t", span, under);
2145 self.x87_at("fld_t", span, top);
2146
2147 let block = self.at.expect("a block is being filled");
2148 let reg = self.new_reg(result);
2149 // Taken before the instruction is started rather than inside it, since both come from the
2150 // same function being built and only one thing at a time may be adding to it.
2151 let spare = both.then(|| self.out.new_vreg(gpr));
2152 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
2153 let mut build = self.out.build(block, opcode).at(span).def(reg, gpr);
2154 if let Some(spare) = spare {
2155 build = build.def(spare, gpr);
2156 }
2157 build.finish();
2158 Ok(())
2159 }
2160
2161 /// The operands and the one result of an instruction that has exactly one.
2162 fn ends(&self, inst: Inst) -> Result<(&'a [Value], Value), Unsupported> {
2163 let data = &self.source[inst];
2164 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
2165 Ok((&self.source[data.args], result))
2166 }
2167
2168 /// The operand of a conversion, which is the end of it that is not the `long double`.
2169 fn narrow(&self, inst: Inst) -> Result<Value, Unsupported> {
2170 let args = &self.source[self.source[inst].args];
2171 args.first().copied().ok_or_else(|| self.unsupported(inst))
2172 }
2173
2174 /// One `va_start`, as the fields of the list it was handed.
2175 ///
2176 /// On the four field list, two of them are numbers this already knows, and each costs an
2177 /// instruction to put in a register before it can be stored, because the machine here has no
2178 /// store of an immediate to memory. The other two are addresses in the frame, and each is a
2179 /// `lea` [`crate::finish`] finishes: the save area is one of the function's own stack objects,
2180 /// and the caller's argument area is where the parameters that had no register came from, which
2181 /// is the same place and the same fixup a parameter past the sixth already uses.
2182 ///
2183 /// On the list that is a pointer it is the second of those four and nothing else, since the
2184 /// whole of what that list says is where the walk is and the walk starts at the first argument
2185 /// the signature does not name. One `lea` and one store.
2186 ///
2187 /// What is written is exactly the fields [`crate::varargs`] describes, in the order they are
2188 /// laid out, so that reading this beside that table is the whole of the check.
2189 fn va_start(&mut self, inst: Inst) -> Result<(), Unsupported> {
2190 let Some(&list) = self.source[self.source[inst].args].first() else {
2191 return Err(self.unsupported(inst));
2192 };
2193 let started = self.varargs.ok_or_else(|| self.unsupported(inst))?;
2194 let list = self.reg_of(list)?;
2195 let block = self.at.expect("a block is being filled");
2196 let span = self.source.span(inst);
2197
2198 let (save, incoming) = match started {
2199 Varargs::Pointer { incoming } => (None, incoming),
2200 Varargs::Fields { save, incoming, integers, floats } => {
2201 for (at, count) in [(varargs::GP_OFFSET, integers), (varargs::FP_OFFSET, floats)] {
2202 let held = self.out.new_vreg(self.gpr);
2203 let load = mir::Opcode::new(self.names.intern("x64.mov_ri_32"));
2204 let build = self.out.build(block, load).at(span);
2205 build.def(held, self.gpr).imm(i64::from(count)).finish();
2206
2207 let store = mir::Opcode::new(self.names.intern("x64.mov_mr_32"));
2208 let mem = self.field(list, at);
2209 self.out.build(block, store).at(span).uses(held, self.gpr).mem(mem).finish();
2210 }
2211 (Some(save), incoming)
2212 }
2213 };
2214
2215 // The first argument the signature did not name, which is as far up the caller's argument
2216 // area as the ones it did name reached. Nothing here knows where that area is, so the
2217 // distance is recorded the way a parameter read out of it is and finished with it.
2218 let overflow = self.out.new_vreg(self.gpr);
2219 let lea = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.lea)));
2220 let sp = mir::Operand::read(mir::Reg::physical(self.conv.stack_pointer), self.gpr);
2221 let made = self
2222 .out
2223 .build(block, lea)
2224 .at(span)
2225 .def(overflow, self.gpr)
2226 .mem(mir::Mem::at(sp))
2227 .finish();
2228 self.stack.arguments.push((made, incoming));
2229
2230 // At the front of the list when that address is the whole of it, and at the field the
2231 // layout gives it when there are four, with the save area behind it.
2232 let fields = match save {
2233 None => vec![(0, overflow)],
2234 Some(save) => {
2235 let save = self.frame_address(block, save);
2236 vec![(varargs::OVERFLOW, overflow), (varargs::SAVE_AREA, save)]
2237 }
2238 };
2239 for (at, held) in fields {
2240 let store = mir::Opcode::new(self.names.intern("x64.mov_mr_64"));
2241 let mem = self.field(list, at);
2242 self.out.build(block, store).at(span).uses(held, self.gpr).mem(mem).finish();
2243 }
2244 Ok(())
2245 }
2246
2247 /// One field of a list, as the addressing mode that reaches it.
2248 fn field(&self, list: mir::Reg, at: i64) -> mir::Mem {
2249 let base = mir::Operand::read(list, self.gpr);
2250 mir::Mem::at(base).plus(i32::try_from(at).expect("a field of a list is a small offset"))
2251 }
2252
2253 /// The address of a name: one `lea` off the instruction pointer, with the name on it.
2254 ///
2255 /// The same instruction an `alloca` gets and for a related reason. An address that is not in
2256 /// the program is a `lea` of an addressing mode that names no register, and the mode carries
2257 /// the symbol so that [`rucc_asm`] can write it relative to `%rip` and leave the relocation
2258 /// for the assembler. Both halves of that already existed: the printer writes `sym(%rip)` and
2259 /// the encoder emits the relocation, because a call to a name the file does not define needed
2260 /// them first.
2261 ///
2262 /// One `mov` and not one `lea` when the name is one [`Elsewhere`] holds, because the distance
2263 /// the `lea` adds to the instruction pointer is a number only a link that puts the name in
2264 /// this program can work out, and the address of a function this file merely declares is not
2265 /// such a number. The load reads the address out of the slot the linker fills in instead. The
2266 /// linker turns it back into the `lea` when the name turns out to have been here all along,
2267 /// so this is not slower in the case that was already right.
2268 ///
2269 /// There is deliberately no name for this in [`crate::term`], which is what stops the address
2270 /// being folded into the instruction that reads it. Folding it is the right thing to do and
2271 /// is what turns a load of a global from two instructions into one, but it is a separate
2272 /// question about addressing modes and issue #282 is it. Until then the address is in a
2273 /// register before anything uses it, which is correct and one instruction longer.
2274 ///
2275 /// What this does not do is give the name anything to refer to. A module carries its globals
2276 /// and nothing writes them out, so a file that defines the variable it reads compiles to a
2277 /// reference the linker cannot resolve. Issue #293 is the other half.
2278 ///
2279 /// A thread-local variable is neither of the two above and is [`Self::thread_address`].
2280 fn address_of(&mut self, inst: Inst) -> Result<(), Unsupported> {
2281 let data = &self.source[inst];
2282 let Extra::Symbol(symbol) = data.extra else { return Err(self.unsupported(inst)) };
2283 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
2284 if self.elsewhere.thread(symbol) {
2285 return self.thread_address(inst, symbol, result);
2286 }
2287
2288 let block = self.at.expect("a block is being filled");
2289 let reg = self.new_reg(result);
2290 let span = self.source.span(inst);
2291 let (mnemonic, mem) = if self.elsewhere.holds(symbol) {
2292 (GOT_LOAD, mir::Mem::got(symbol))
2293 } else {
2294 (x86_64::FRAME.lea, mir::Mem::of(symbol))
2295 };
2296 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{mnemonic}")));
2297 self.out.build(block, opcode).at(span).def(reg, self.gpr).mem(mem).finish();
2298 Ok(())
2299 }
2300
2301 /// The address of a thread-local variable, which is this thread's copy of it.
2302 ///
2303 /// Neither instruction the ordinary case writes would mean anything here. There is no distance
2304 /// to the variable for a `lea` to add, because there is no variable: there is one copy of it per
2305 /// thread and they are at different addresses, so a link asked for the distance to the name
2306 /// refuses rather than picking one. And there is no address for a table slot to hold either, for
2307 /// the same reason.
2308 ///
2309 /// What is the same in every thread is where the variable sits inside the block of storage a
2310 /// thread gets, so that offset is what the link writes down, and the address of the running
2311 /// thread's block is what turns it into an address. x86-64 keeps that address in `%fs`, at the
2312 /// front of the block, so the whole of this is three instructions:
2313 ///
2314 /// ```text
2315 /// movq x@gottpoff(%rip), %off # how far into the block x sits, which the link fills in
2316 /// movq %fs:0, %tp # where this thread's block is, which only the machine knows
2317 /// addq %tp, %off # this thread's copy of x
2318 /// ```
2319 ///
2320 /// That is the initial exec model. It is one instruction longer than what gcc writes at `-O2`
2321 /// in an executable, which folds the addition into the instruction that uses the address, and
2322 /// the difference is issue #282 rather than anything about threads: nothing here folds an
2323 /// address into its reader yet. The link relaxes the first instruction into an immediate when it
2324 /// is making an executable, since it lays the blocks out and therefore knows the number, so the
2325 /// table slot costs nothing in the case that is common.
2326 ///
2327 /// It is not the most general model. A library loaded by `dlopen` gets its storage after the
2328 /// program is already running, and the block this reaches was laid out before it started, so
2329 /// the loader has to find room in that block for the library's variables. glibc keeps a little
2330 /// spare room for exactly this and a library that fits in it loads and runs; one that does not
2331 /// fails to load, with a message saying so. The model with no such limit calls `__tls_get_addr`
2332 /// and is what gcc writes under `-fPIC` by default, and it is issue #1104.
2333 ///
2334 /// So this is the model gcc writes under `-ftls-model=initial-exec`: right for an executable,
2335 /// right for a library the program is linked against, and a load that either works or is
2336 /// refused out loud for a library something opens later. What it is never is quietly wrong.
2337 fn thread_address(
2338 &mut self,
2339 inst: Inst,
2340 symbol: Symbol,
2341 result: Value,
2342 ) -> Result<(), Unsupported> {
2343 let block = self.at.expect("a block is being filled");
2344 let span = self.source.span(inst);
2345 let gpr = self.gpr;
2346 let load = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{GOT_LOAD}")));
2347
2348 let offset = self.out.new_vreg(gpr);
2349 self.out
2350 .build(block, load)
2351 .at(span)
2352 .def(offset, gpr)
2353 .mem(mir::Mem::thread(symbol))
2354 .finish();
2355 // The front of the block, which is the one thing on this machine that no instruction can
2356 // work out: `%fs` is not a register a program can read, and what it points at is a word
2357 // holding its own address, so reading through it at zero is how the address is come by.
2358 let pointer = self.out.new_vreg(gpr);
2359 let at = mir::Mem::in_segment(Segment::Fs, 0);
2360 self.out.build(block, load).at(span).def(pointer, gpr).mem(at).finish();
2361
2362 // Two address, spelled out for the reason `x87_to_int` gives: this machine adds into the
2363 // register it read, and only the constraint says the two are the same one.
2364 let reg = self.new_reg(result);
2365 let add = mir::Opcode::new(self.names.intern(&format!("{PREFIX}add_rr_64")));
2366 self.out
2367 .build(block, add)
2368 .at(span)
2369 .operand(mir::Operand::write(reg, gpr).with(Constraint::Reuse(1)))
2370 .operand(mir::Operand::read(offset, gpr))
2371 .operand(mir::Operand::read(pointer, gpr))
2372 .finish();
2373 Ok(())
2374 }
2375
2376 /// `&&label`, GNU's address of a label, which is the same `lea` a global gets against a place
2377 /// in this same function.
2378 ///
2379 /// What the two have in common is the whole of the instruction: an address worked out from
2380 /// where the instruction is, which is what `(%rip)` means and is the only way this compiler
2381 /// reaches anything. What they do not have in common is what fills the four bytes in. A
2382 /// global is a name, so the number is a relocation and the linker writes it. A block is a
2383 /// place in this function, so both ends are in one section and the number is known as soon as
2384 /// the blocks have been laid out, which is why `rucc_asm` fills it in the way it fills in a
2385 /// jump rather than leaving a relocation behind.
2386 ///
2387 /// Nothing here says the block is one control can arrive at. That is said by the
2388 /// [`Opcode::IndirectBr`] that reads the address, which lists every block it can arrive at,
2389 /// and by nothing else: an address on its own is a number.
2390 fn block_address(&mut self, inst: Inst) -> Result<(), Unsupported> {
2391 let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
2392 let Some(call) = self.source.successors(inst).next() else {
2393 return Err(self.unsupported(inst));
2394 };
2395 let block = self.at.expect("a block is being filled");
2396 let reg = self.new_reg(result);
2397 let span = self.source.span(inst);
2398 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.lea)));
2399 let mem = mir::Mem::block(self.out_block(call.block));
2400 self.out.build(block, opcode).at(span).def(reg, self.gpr).mem(mem).finish();
2401 Ok(())
2402 }
2403
2404 /// `goto *p`, GNU's computed goto, which is a jump through a register.
2405 ///
2406 /// Where it goes is not written here and cannot be. Every block it can arrive at is on the
2407 /// block this ends, the way every other arm is, and which of them the address holds is decided
2408 /// while the program runs. So this is one instruction with one operand, and the arms are
2409 /// copied across by [`Self::edges`] like anybody else's.
2410 fn indirect_branch(&mut self, inst: Inst) -> Result<(), Unsupported> {
2411 let data = &self.source[inst];
2412 let &address = self.source[data.args].first().ok_or_else(|| self.unsupported(inst))?;
2413 let reg = self.reg_of(address)?;
2414 let block = self.at.expect("a block is being filled");
2415 let span = self.source.span(inst);
2416 let name = x86_64::BRANCH.indirect;
2417 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
2418 self.out.build(block, opcode).at(span).operand(mir::Operand::read(reg, self.gpr)).finish();
2419 Ok(())
2420 }
2421
2422 /// A `switch` on an index from zero up, as a jump through a table of this function.
2423 ///
2424 /// Every `switch` that reaches here is one `crate::switch` left behind on purpose: it has
2425 /// already checked the value is inside the table and taken the lowest case off it, so the
2426 /// operand is a 64 bit index, the cases are the values from zero up with gaps where the
2427 /// program had no case, and the default is only where those gaps go. What is written is the
2428 /// shape gcc writes for the same statement in position independent code:
2429 ///
2430 /// ```text
2431 /// leaq table(%rip), %base
2432 /// movslq (%base,%index,4), %offset
2433 /// addq %base, %offset
2434 /// jmp *%offset
2435 /// ```
2436 ///
2437 /// The table holds distances from itself to each arm rather than addresses, which is what
2438 /// lets it be filled in by the assembler with nothing left for a linker to do. Each cell is
2439 /// stored as the place of an arm among this block's successors, which [`Self::edges`] copies
2440 /// across in the IR's own order, the default first and then one per case. See
2441 /// [`mir::Table`] for why a place and not a block.
2442 fn jump_table(&mut self, inst: Inst) -> Result<(), Unsupported> {
2443 let data = &self.source[inst];
2444 let Extra::Switch(info) = data.extra else { return Err(self.unsupported(inst)) };
2445 let &index = self.source[data.args].first().ok_or_else(|| self.unsupported(inst))?;
2446 let ty = self.source[index].ty;
2447 if ty != Type::int(u64::BITS) {
2448 return Err(self.unsupported(inst));
2449 }
2450 let cases = self.source[self.source[info].cases].to_vec();
2451 let mut cells: Vec<u32> = Vec::new();
2452 for (arm, case) in cases.iter().enumerate() {
2453 let at = usize::try_from(case.signed(ty)).map_err(|_| self.unsupported(inst))?;
2454 if at >= cells.len() {
2455 cells.resize(at + 1, 0);
2456 }
2457 cells[at] = u32::try_from(arm + 1).map_err(|_| self.unsupported(inst))?;
2458 }
2459 let reg = self.reg_of(index)?;
2460 let block = self.at.expect("a block is being filled");
2461 let span = self.source.span(inst);
2462 let gpr = self.gpr;
2463 let table = u32::try_from(self.out.tables.len()).expect("fewer tables than that");
2464
2465 let base = self.out.new_vreg(gpr);
2466 let lea = self.named(x86_64::FRAME.lea);
2467 self.out.build(block, lea).at(span).def(base, gpr).mem(mir::Mem::table(table)).finish();
2468 let offset = self.out.new_vreg(gpr);
2469 let cell =
2470 mir::Mem::at(mir::Operand::read(base, gpr)).indexed(mir::Operand::read(reg, gpr), 4);
2471 let load = self.named("movsxd_rm_32_64");
2472 self.out.build(block, load).at(span).def(offset, gpr).mem(cell).finish();
2473 // Two address, for the reason `thread_pointer` gives.
2474 let to = self.out.new_vreg(gpr);
2475 let add = self.named("add_rr_64");
2476 self.out
2477 .build(block, add)
2478 .at(span)
2479 .operand(mir::Operand::write(to, gpr).with(Constraint::Reuse(1)))
2480 .operand(mir::Operand::read(offset, gpr))
2481 .operand(mir::Operand::read(base, gpr))
2482 .finish();
2483 let jump = self.named(x86_64::BRANCH.indirect);
2484 let jump =
2485 self.out.build(block, jump).at(span).operand(mir::Operand::read(to, gpr)).finish();
2486 self.out.tables.push(mir::Table { jump, cells });
2487 Ok(())
2488 }
2489
2490 /// `__builtin_setjmp`, which writes down where the function is so that a `__builtin_longjmp`
2491 /// somewhere else can bring control back here, and answers zero on the way past.
2492 ///
2493 /// Four words of the buffer, the three gcc writes and one of this compiler's own, and then the
2494 /// block ends: everything after the save in the IR block is put into a new machine IR block,
2495 /// and the address of that block is what went into the buffer. That is the whole reason the
2496 /// block is split here. An address points at a label, a machine IR block is the only thing in
2497 /// this representation that has one, and a save is in the middle of a block rather than at the
2498 /// end of one.
2499 ///
2500 /// # How the answer gets back
2501 ///
2502 /// Through the frame rather than through a register. The save writes a zero into a word of its
2503 /// own frame, puts the address of that word in the buffer, and the new block reads the word
2504 /// back. The restore writes a one through the address it finds in the buffer before it goes.
2505 /// So one load answers zero on the way past and one on the way back, and neither path has to
2506 /// agree with the other about a register.
2507 ///
2508 /// gcc does it the other way round, with a second block that sets the answer to one and is
2509 /// what the restore arrives at. That block is one nothing in the function jumps to, and a
2510 /// machine IR whose blocks are walked from the entry has nowhere to put such a thing: the
2511 /// allocator lays a function out in the line it is going to be emitted in, and a block no edge
2512 /// reaches is not in that line. The word in the frame costs eight bytes of stack and one load,
2513 /// and it needs nothing said anywhere about a block arrived at from outside.
2514 ///
2515 /// # What the allocator is told
2516 ///
2517 /// That every register it hands out is gone at the end of the first block. That is what makes
2518 /// the rest of the function right on the way back: control arrives from a `__builtin_longjmp`
2519 /// in some other function, and the only two registers that puts back are the stack pointer and
2520 /// the frame pointer, so anything this function still wants has to be in the frame those two
2521 /// reach. It is said with a write of every one of those registers, which is the same thing a
2522 /// call says about the registers a callee may destroy, on an instruction with nothing else on
2523 /// it so that the stores above are not caught up in it.
2524 fn saves_place(&mut self, inst: Inst) -> Result<(), Unsupported> {
2525 let data = &self.source[inst];
2526 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
2527 let &buffer = self.source[data.args].first().ok_or_else(|| self.unsupported(inst))?;
2528 let span = self.source.span(inst);
2529 let buf = self.reg_of(buffer)?;
2530 let at = self.at.expect("a block is being filled");
2531 let gpr = self.gpr;
2532 let moves = x86_64::FRAME.moves(gpr).expect("a class the target says how to move");
2533 let store = self.named(moves.store);
2534 let load = self.named(moves.load);
2535 let lea = self.named(x86_64::FRAME.lea);
2536 let put = self.named(x86_64::FRAME.imm);
2537 let nothing = x86_64::FRAME.pad.expect("a target with an instruction that does nothing");
2538 let nothing = self.named(nothing);
2539 self.stack.saves_place = true;
2540 let answer = self.answer_slot();
2541 let back = self.out.create_block();
2542
2543 // The zero this answers with, into the word a restore writes a one into.
2544 let zero = self.out.new_vreg(gpr);
2545 self.out.build(at, put).at(span).def(zero, gpr).imm(0).finish();
2546 let mem = self.frame_mem();
2547 let made = self.out.build(at, store).at(span).uses(zero, gpr).mem(mem).finish();
2548 self.stack.addresses.push((made, answer));
2549
2550 // The four words: where that word is, where control comes back to, and the two registers
2551 // the restore puts back.
2552 let found = self.frame_address(at, answer);
2553 self.write_word(at, span, store, found, buf, JUMP_ANSWER);
2554 let pc = self.out.new_vreg(gpr);
2555 self.out.build(at, lea).at(span).def(pc, gpr).mem(mir::Mem::block(back)).finish();
2556 self.write_word(at, span, store, pc, buf, JUMP_PC);
2557 let frame = mir::Reg::physical(self.conv.frame_pointer);
2558 self.write_word(at, span, store, frame, buf, JUMP_FRAME);
2559 let stack = mir::Reg::physical(self.conv.stack_pointer);
2560 self.write_word(at, span, store, stack, buf, JUMP_STACK);
2561
2562 // Nothing is in a register past this point, which is what the rest of the function is
2563 // allowed to assume about the way back in.
2564 let gone = self.across_jump();
2565 let mut build = self.out.build(at, nothing).at(span);
2566 for (reg, class) in gone {
2567 build = build.operand(mir::Operand::write(reg, class));
2568 }
2569 build.finish();
2570
2571 // And the rest of the block, which is the block the address above was of.
2572 *self.out.succs_mut(at) = vec![mir::BlockCall::to(back)];
2573 self.at = Some(back);
2574 let reg = self.new_reg(result);
2575 let mem = self.frame_mem();
2576 let made = self.out.build(back, load).at(span).def(reg, gpr).mem(mem).finish();
2577 self.stack.addresses.push((made, answer));
2578 Ok(())
2579 }
2580
2581 /// `__builtin_longjmp`, which reads a buffer a `__builtin_setjmp` filled in and goes there.
2582 ///
2583 /// Everything comes out of the buffer before anything is put back, and the four registers it
2584 /// comes out into are physical ones rather than values the allocator places. Both of those are
2585 /// about the same moment. The stack pointer is one of the things being put back, a value the
2586 /// allocator sent to the stack is reached through the stack pointer, and between the
2587 /// instruction that moves it and the jump there is no stack this function owns any more. A
2588 /// register named outright is a register nothing reloads into and nothing else is in, which is
2589 /// the only way to hold something across that moment.
2590 ///
2591 /// Four of them because that is how many things are in the air at once: where to go, the frame
2592 /// pointer to put back, the one the matching save is to answer with, and one register used
2593 /// twice, first for the address that one is written through and then for the stack pointer.
2594 ///
2595 /// Nothing after this in the block is reached. The marker is not a terminator, for the reason
2596 /// `spec/08-ir.md` gives, so the block goes on and whatever the front end wrote after it is
2597 /// written out and never run.
2598 fn comes_back(&mut self, inst: Inst) -> Result<(), Unsupported> {
2599 let data = &self.source[inst];
2600 let &buffer = self.source[data.args].first().ok_or_else(|| self.unsupported(inst))?;
2601 let span = self.source.span(inst);
2602 let buf = self.reg_of(buffer)?;
2603 let at = self.at.expect("a block is being filled");
2604 let gpr = self.gpr;
2605 let moves = x86_64::FRAME.moves(gpr).expect("a class the target says how to move");
2606 let load = self.named(moves.load);
2607 let store = self.named(moves.store);
2608 let mov = self.named(moves.mov);
2609 let put = self.named(x86_64::FRAME.imm);
2610 let jump = self.named(x86_64::BRANCH.indirect);
2611
2612 let held = self.jump_regs();
2613 if held.len() < JUMP_REGS {
2614 return Err(self.unsupported(inst));
2615 }
2616 let pc = mir::Reg::physical(held[0]);
2617 let frame = mir::Reg::physical(held[1]);
2618 let spare = mir::Reg::physical(held[2]);
2619 let one = mir::Reg::physical(held[3]);
2620
2621 self.read_word(at, span, load, pc, buf, JUMP_PC);
2622 self.read_word(at, span, load, frame, buf, JUMP_FRAME);
2623 self.read_word(at, span, load, spare, buf, JUMP_ANSWER);
2624
2625 // What the matching save answers with, written through the address that came out of the
2626 // buffer, because the word it goes in is in the other function's frame and this one has no
2627 // way of knowing where that is.
2628 self.out.build(at, put).at(span).def(one, gpr).imm(1).finish();
2629 let mem = mir::Mem::at(mir::Operand::read(spare, gpr));
2630 self.out.build(at, store).at(span).uses(one, gpr).mem(mem).finish();
2631
2632 // The stack last of the four, so that the register the buffer is reached through is done
2633 // with before the stack it may have been spilled to stops being this function's.
2634 self.read_word(at, span, load, spare, buf, JUMP_STACK);
2635 let stack = mir::Reg::physical(self.conv.stack_pointer);
2636 self.copy(at, span, mov, stack, spare);
2637 let base = mir::Reg::physical(self.conv.frame_pointer);
2638 self.copy(at, span, mov, base, frame);
2639
2640 // And the jump, which reads the two registers just put back as well as the address it
2641 // goes through. Neither of those is printed, because the target's spelling of an indirect
2642 // jump has one argument and it is the first one read. They are there because the code
2643 // control arrives at reaches its frame through them, and because without them the two
2644 // instructions above write registers nothing reads: a scheduler is then free to put the
2645 // jump in front of them, and at `-O2` it does.
2646 self.out
2647 .build(at, jump)
2648 .at(span)
2649 .operand(mir::Operand::read(pc, gpr))
2650 .operand(mir::Operand::read(stack, gpr))
2651 .operand(mir::Operand::read(base, gpr))
2652 .finish();
2653 Ok(())
2654 }
2655
2656 /// One word of the buffer of a `__builtin_setjmp`, written from a register.
2657 fn write_word(
2658 &mut self,
2659 at: mir::Block,
2660 span: Span,
2661 store: mir::Opcode,
2662 from: mir::Reg,
2663 buf: mir::Reg,
2664 word: i32,
2665 ) {
2666 let mem = mir::Mem::at(mir::Operand::read(buf, self.gpr)).plus(word);
2667 self.out.build(at, store).at(span).uses(from, self.gpr).mem(mem).finish();
2668 }
2669
2670 /// One word of that buffer, read back into a register.
2671 fn read_word(
2672 &mut self,
2673 at: mir::Block,
2674 span: Span,
2675 load: mir::Opcode,
2676 into: mir::Reg,
2677 buf: mir::Reg,
2678 word: i32,
2679 ) {
2680 let mem = mir::Mem::at(mir::Operand::read(buf, self.gpr)).plus(word);
2681 self.out.build(at, load).at(span).def(into, self.gpr).mem(mem).finish();
2682 }
2683
2684 /// One register into another, which is the one shape of instruction the builder has no word
2685 /// for because neither operand is a definition of a value or a read of memory.
2686 fn copy(
2687 &mut self,
2688 at: mir::Block,
2689 span: Span,
2690 mov: mir::Opcode,
2691 into: mir::Reg,
2692 from: mir::Reg,
2693 ) {
2694 self.out
2695 .build(at, mov)
2696 .at(span)
2697 .operand(mir::Operand::write(into, self.gpr))
2698 .operand(mir::Operand::read(from, self.gpr))
2699 .finish();
2700 }
2701
2702 /// The word a `__builtin_setjmp` in this function answers with, asked for once and kept.
2703 fn answer_slot(&mut self) -> usize {
2704 match self.answer {
2705 Some(index) => index,
2706 None => {
2707 let index = self.stack.locals.len();
2708 self.stack.locals.push(Local { size: JUMP_WORD, align: JUMP_WORD });
2709 self.answer = Some(index);
2710 index
2711 }
2712 }
2713 }
2714
2715 /// An address in this function's frame with nothing in its displacement, which is what an
2716 /// instruction reaching one of its stack objects is written with until [`crate::finish`] knows
2717 /// where the object is.
2718 fn frame_mem(&self) -> mir::Mem {
2719 mir::Mem::at(mir::Operand::read(mir::Reg::physical(self.conv.stack_pointer), self.gpr))
2720 }
2721
2722 /// Every register the allocator hands out, which is what a `__builtin_setjmp` destroys.
2723 ///
2724 /// Both files, since a `double` live across a save has the same problem an integer does. The
2725 /// two registers a frame is reached through are not here: the restore puts both of them back,
2726 /// which is the whole of what it puts back, and a function whose frame pointer was destroyed
2727 /// by its own save would have nothing left to find its caller with.
2728 fn across_jump(&self) -> Vec<(mir::Reg, RegClass)> {
2729 let mut gone = Vec::new();
2730 for ® in self.conv.int_order {
2731 if reg == self.conv.stack_pointer || reg == self.conv.frame_pointer {
2732 continue;
2733 }
2734 gone.push((mir::Reg::physical(reg), self.gpr));
2735 }
2736 for ® in self.conv.sse_order {
2737 gone.push((mir::Reg::physical(reg), self.conv.sse_class));
2738 }
2739 gone
2740 }
2741
2742 /// The registers a `__builtin_longjmp` may hold things in while it puts a frame back.
2743 ///
2744 /// The ones the allocator hands out, less the two a frame is reached through. The scratch
2745 /// registers are not among them on purpose: the rewriter writes a reload into one of those
2746 /// wherever it likes, and one of these has to survive from the load that fills it to the
2747 /// instruction that reads it however many instructions apart those are.
2748 fn jump_regs(&self) -> Vec<PhysReg> {
2749 self.conv
2750 .int_order
2751 .iter()
2752 .copied()
2753 .filter(|®| {
2754 reg != self.conv.stack_pointer
2755 && reg != self.conv.frame_pointer
2756 && !crate::pipeline::SCRATCH.contains(®)
2757 })
2758 .collect()
2759 }
2760
2761 /// A machine opcode of this target from the name the target gives it.
2762 fn named(&mut self, name: &str) -> mir::Opcode {
2763 mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")))
2764 }
2765
2766 /// `__builtin_frame_address` and `__builtin_return_address`, which are a walk up the chain of
2767 /// saved frame pointers and then one thing read at the end of it.
2768 ///
2769 /// Every frame that kept a frame pointer holds the caller's at the address the register points
2770 /// at, and the address that frame returns to one word above that, which is where the call
2771 /// instruction put it and where the prologue's push left it. So the walk is a load through the
2772 /// register for each link, the frame address is wherever the walk stopped, and the return
2773 /// address is one more load from a word above it. gcc 16.2.0 writes exactly this, measured on
2774 /// x86-64 at `-O2` for depths zero to three of both builtins.
2775 ///
2776 /// The function is given a frame pointer because of this, which is what [`Stack::walks_frames`]
2777 /// carries out to the layout. A depth of zero needs it as the answer and every depth above zero
2778 /// needs it as the start, so there is no case here where it is not wanted.
2779 ///
2780 /// How far the chain actually reaches is the program's business and not this one's. A caller
2781 /// compiled without a frame pointer has no link in it for the walk to follow, so a depth above
2782 /// zero is a promise about how the whole program was built. That is why gcc documents a nonzero
2783 /// depth as unsafe rather than as an answer, and why the depth is refused above a limit in
2784 /// `check/builtin/frame.rs` rather than walked as far as it says.
2785 fn frames(&mut self, inst: Inst) -> Result<(), Unsupported> {
2786 let data = &self.source[inst];
2787 let Extra::Depth(depth) = data.extra else { return Err(self.unsupported(inst)) };
2788 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
2789 let returning = data.opcode == Opcode::ReturnAddress;
2790 let block = self.at.expect("a block is being filled");
2791 let span = self.source.span(inst);
2792 let moves = x86_64::FRAME.moves(self.gpr).expect("a class the target says how to move");
2793 let load = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", moves.load)));
2794 self.stack.walks_frames = true;
2795
2796 // Where the walk is up to. The frame pointer to begin with, and the register the last load
2797 // wrote after that.
2798 let reg = self.new_reg(result);
2799 let mut base = mir::Reg::physical(self.conv.frame_pointer);
2800 for link in 0..depth {
2801 // The last load of a walk that is looking for a frame writes the answer itself, which
2802 // is what keeps a walk of so many links that many instructions and not one more.
2803 let ends_here = link + 1 == depth && !returning;
2804 let next = if ends_here { reg } else { self.out.new_vreg(self.gpr) };
2805 let at = mir::Mem::at(mir::Operand::read(base, self.gpr));
2806 self.out.build(block, load).at(span).def(next, self.gpr).mem(at).finish();
2807 base = next;
2808 }
2809
2810 if returning {
2811 let up = i32::try_from(self.conv.return_address).expect("a word above the frame");
2812 let at = mir::Mem::at(mir::Operand::read(base, self.gpr)).plus(up);
2813 self.out.build(block, load).at(span).def(reg, self.gpr).mem(at).finish();
2814 } else if depth == 0 {
2815 // The one case with no load in it at all: the frame this function is running in is the
2816 // register itself, and a physical register is not one the allocator hands out, so the
2817 // answer is a copy of it.
2818 let mov = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", moves.mov)));
2819 self.out
2820 .build(block, mov)
2821 .at(span)
2822 .operand(mir::Operand::write(reg, self.gpr))
2823 .operand(mir::Operand::read(base, self.gpr))
2824 .finish();
2825 }
2826 Ok(())
2827 }
2828
2829 /// `__builtin_thread_pointer`, which is the front of the block [`Self::thread_address`] adds
2830 /// an offset to.
2831 ///
2832 /// The same one instruction, on its own this time and with nothing to add to it. A program
2833 /// writes this when what it wants is a number that is different in every thread and cheap to
2834 /// come by, rather than a variable of its own in the block, so there is no relocation here and
2835 /// no name for the link to resolve.
2836 fn thread_pointer(&mut self, inst: Inst) -> Result<(), Unsupported> {
2837 let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
2838 let block = self.at.expect("a block is being filled");
2839 let span = self.source.span(inst);
2840 let reg = self.new_reg(result);
2841 let load = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{GOT_LOAD}")));
2842 let at = mir::Mem::in_segment(Segment::Fs, 0);
2843 self.out.build(block, load).at(span).def(reg, self.gpr).mem(at).finish();
2844 Ok(())
2845 }
2846
2847 /// What a named machine register holds, which is `register long x asm ("rbx");`.
2848 ///
2849 /// One move out of that register, with the register named as itself the way a register a
2850 /// template wrote is named, which is [`Self::itself`] and is the thing #1653 built. What it
2851 /// buys here is what it buys there: the register is part of the instruction the allocator
2852 /// sees, so it is a use the allocator will not have written over first, and the value goes
2853 /// into an ordinary one of its own that everything downstream reads.
2854 ///
2855 /// The whole sixty four bits are moved whatever the type is, because the register is that
2856 /// wide and a narrower type reads the low end of the copy, which is the same low end. A type
2857 /// wider than the register is refused, since there is no register holding it to read.
2858 ///
2859 /// A name the machine has not got is refused too, and is the only thing that can be wrong
2860 /// with the string: which register a name means is this machine's question and this is where
2861 /// the question is asked, at the same table `asm` asks about clobbers at. The sigil gcc
2862 /// allows in front of it is taken off here, because what the name is written with is syntax.
2863 fn register_value(&mut self, inst: Inst) -> Result<(), Unsupported> {
2864 let Extra::Symbol(symbol) = self.source[inst].extra else {
2865 return Err(self.unsupported(inst));
2866 };
2867 let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
2868 let ty = self.source[result].ty;
2869 let bits = if ty.is_ptr() { ADDRESS_BITS } else { ty.bits() };
2870 if bits > ADDRESS_BITS {
2871 return Err(self.unsupported(inst));
2872 }
2873 let spelled = self.names.resolve(symbol).to_owned();
2874 let named = x86_64::gpr_named(spelled.strip_prefix('%').unwrap_or(&spelled));
2875 let Some((held, _)) = named else {
2876 return Err(Unsupported::Register { inst, name: spelled });
2877 };
2878 let block = self.at.expect("a block is being filled");
2879 let span = self.source.span(inst);
2880 let mov = x86_64::FRAME.moves(self.gpr).expect("a class the target says how to move").mov;
2881 let mov = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{mov}")));
2882 let into = self.new_reg(result);
2883 self.out
2884 .build(block, mov)
2885 .at(span)
2886 .operand(mir::Operand::write(into, self.gpr))
2887 .operand(
2888 mir::Operand::read(mir::Reg::physical(held), self.gpr)
2889 .with(Constraint::Fixed(held)),
2890 )
2891 .finish();
2892 Ok(())
2893 }
2894
2895 /// A conversion that converts nothing: the result is the operand under another type.
2896 ///
2897 /// `ptrtoint` and `inttoptr` at one width are the whole of this. An address on this machine is
2898 /// an integer as wide as the machine addresses, so a cast between the two changes what the
2899 /// type system calls the value and changes nothing about the value, and the register holding
2900 /// it is the register that already held it. The front end never writes either of them at any
2901 /// other width, because it widens or narrows around the cast rather than through it, so the
2902 /// two widths disagreeing here means the IR came from somewhere else and is refused rather
2903 /// than guessed at.
2904 ///
2905 /// Reading the operand first is what materializes it when it is a constant, which is the case
2906 /// that matters: a null pointer is an `inttoptr` of zero, and that zero has to reach a
2907 /// register before anything can call it an address.
2908 fn rename(&mut self, inst: Inst) -> Result<(), Unsupported> {
2909 let data = &self.source[inst];
2910 let [arg] = self.source[data.args] else { return Err(self.unsupported(inst)) };
2911 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
2912 if !self.is_address_width(self.source[arg].ty)
2913 || !self.is_address_width(self.source[result].ty)
2914 {
2915 return Err(self.unsupported(inst));
2916 }
2917 let reg = self.reg_of(arg)?;
2918 self.regs[result.index()] = Some(reg);
2919 Ok(())
2920 }
2921
2922 /// One barrier, which on this machine is one instruction at the strongest ordering and no
2923 /// instruction at all at every other one.
2924 ///
2925 /// x86-64 is total store order, so the only reordering the machine does is a store followed by
2926 /// a load of a different address, and the only ordering that forbids that is sequential
2927 /// consistency. An acquire, a release and an acquire release fence are therefore already true
2928 /// of every program running here, and what a program wanted from writing one is that the
2929 /// compiler not move memory accesses across it. The optimizer has finished by the time this
2930 /// runs and nothing below reorders one access past another, so the constraint is already
2931 /// discharged and there is nothing to write.
2932 ///
2933 /// The strongest one is `mfence`, which is what gcc 16.2.0 writes for
2934 /// `__atomic_thread_fence(__ATOMIC_SEQ_CST)` and for `__sync_synchronize`. A locked instruction
2935 /// on the stack is faster on most parts and is what some compilers write instead; it is also a
2936 /// write to memory the program did not ask for, and the plain barrier is the one that says what
2937 /// it means.
2938 ///
2939 /// Written here by name rather than by a rule, for the same reason a `lea` of a symbol is:
2940 /// there is nothing in a barrier that a proof over bitvectors could discharge. It computes
2941 /// nothing, so there is no equality to state, and what makes it the right answer is the memory
2942 /// model, which the rule language cannot talk about.
2943 fn barrier(&mut self, inst: Inst) -> Result<(), Unsupported> {
2944 let Extra::Order(order) = self.source[inst].extra else {
2945 return Err(self.unsupported(inst));
2946 };
2947 if order != MemOrder::SeqCst {
2948 return Ok(());
2949 }
2950 let block = self.at.expect("a block is being filled");
2951 let span = self.source.span(inst);
2952 let fence = mir::Opcode::new(self.names.intern("x64.mfence"));
2953 self.out.build(block, fence).at(span).finish();
2954 Ok(())
2955 }
2956
2957 /// The instruction a program stops on, which is one byte pair and no operands.
2958 ///
2959 /// `ud2` is an opcode the manual promises will never be given a meaning, so a processor that
2960 /// reaches it raises the fault for an instruction it does not know, and on Linux that arrives
2961 /// at the program as `SIGILL`. That is what `__builtin_trap` is for: a stop that cannot be
2962 /// caught by anything the program installed for an ordinary error, cannot be returned from,
2963 /// and leaves the address of the fault in the core file.
2964 ///
2965 /// Why not a call to `abort`. It is two bytes against a call and a relocation, it needs no
2966 /// library, and it works in the places this one is written most, which are a kernel and a
2967 /// freestanding program that has no `abort` to call. gcc 16.2.0 writes `ud2` here too.
2968 fn trap(&mut self, inst: Inst) {
2969 let block = self.at.expect("a block is being filled");
2970 let span = self.source.span(inst);
2971 let stop = mir::Opcode::new(self.names.intern("x64.ud2"));
2972 self.out.build(block, stop).at(span).finish();
2973 }
2974
2975 /// One hint that an address is about to be used, which is one instruction and no promise.
2976 ///
2977 /// Four instructions on this machine and the locality picks between them, which is what the
2978 /// number means: how much of the data will still be wanted after the access. None of it wanted
2979 /// is `prefetchnta`, which brings the line in without keeping it, and all of it wanted is
2980 /// `prefetcht0`, which brings it as close as the machine can. The two in between are the levels
2981 /// between those. Measured against gcc 16.2.0 on x86-64 rather than read off the manual: zero
2982 /// gives `prefetchnta`, one `prefetcht2`, two `prefetcht1` and three `prefetcht0`.
2983 ///
2984 /// Whether the access will write is not read here, and that is this machine rather than an
2985 /// omission. The write hint is `prefetchw`, which is not in the base instruction set, and gcc
2986 /// writes it only when the command line said the part has it. So a prefetch for a write is the
2987 /// same instruction as a prefetch for a read, which is what gcc 16.2.0 writes without
2988 /// `-mprfchw`, and the difference is carried in the IR for a target that can use it.
2989 ///
2990 /// The address goes in the addressing mode rather than in an operand, the way a store's does.
2991 /// It is built here as the plainest one there is, a register and nothing else, because what
2992 /// arrives is a value and folding an addition into the mode is a rule's job and no rule reaches
2993 /// this instruction. An address the program computed is therefore one `lea` or one add in front
2994 /// of this, which is what it would have been for the load the hint is about anyway.
2995 fn hint(&mut self, inst: Inst) -> Result<(), Unsupported> {
2996 let Extra::Prefetch(hint) = self.source[inst].extra else {
2997 return Err(self.unsupported(inst));
2998 };
2999 let args: Vec<Value> = self.source[self.source[inst].args].to_vec();
3000 let [address] = args[..] else { return Err(self.unsupported(inst)) };
3001 let name = match hint.locality {
3002 0 => "prefetch_nta",
3003 1 => "prefetch_t2",
3004 2 => "prefetch_t1",
3005 PrefetchHint::MOST => "prefetch_t0",
3006 // Nothing else exists. The checker reads a locality outside the range as zero and the
3007 // verifier refuses one that got here another way, so this is a hint that was built
3008 // rather than checked, and the safe answer for a hint is to write no instruction.
3009 _ => return Err(self.unsupported(inst)),
3010 };
3011 let base = self.reg_of(address)?;
3012 let block = self.at.expect("a block is being filled");
3013 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
3014 self.out
3015 .build(block, opcode)
3016 .at(self.source.span(inst))
3017 .mem(mir::Mem::at(mir::Operand::read(base, self.gpr)))
3018 .finish();
3019 Ok(())
3020 }
3021
3022 /// One compare and exchange, which is the instruction every other atomic on this machine is
3023 /// built out of.
3024 ///
3025 /// What the IR asks for is: read what is at an address, compare it against a value the program
3026 /// expected, put a second value there if the two were equal, and say both what was read and
3027 /// whether the exchange happened. The machine has exactly that instruction, and the `lock` in
3028 /// front of it is what makes the whole of it one step as far as every other processor is
3029 /// concerned.
3030 ///
3031 /// The ordering is not read here, and that is the memory model rather than an omission. A
3032 /// locked instruction on x86-64 is a full barrier whatever the program asked for, so a relaxed
3033 /// compare and exchange and a sequentially consistent one are the same instruction, and there
3034 /// is nothing weaker to emit for the weaker orderings. The failure ordering is not read for the
3035 /// same reason.
3036 ///
3037 /// The two values it produces are why this is written by name. The one the program compares
3038 /// against and the one it gets back are both `rax`, which the instruction reads and writes
3039 /// without being told, and the table says so with a fixed constraint at each end rather than
3040 /// leaving the allocator to find out. The second value is the byte behind it, which is the zero
3041 /// flag read out by a `sete`, and it is a definition of the same instruction so that the
3042 /// allocator knows the two are live together and never gives the byte the register the answer
3043 /// is in.
3044 fn exchange(&mut self, inst: Inst) -> Result<(), Unsupported> {
3045 let args: Vec<Value> = self.source[self.source[inst].args].to_vec();
3046 let results: Vec<Value> = self.source[inst].results().collect();
3047 let [addr, expected, desired] = args[..] else { return Err(self.unsupported(inst)) };
3048 let [old, exchanged] = results[..] else { return Err(self.unsupported(inst)) };
3049
3050 // A value the machine can compare in one instruction, which is an integer or an address at
3051 // one of the four widths it has a compare and exchange for. Anything else is a type this
3052 // has no instruction for rather than a program that is wrong, and the front end refuses it
3053 // before ever getting here.
3054 let ty = self.source[old].ty;
3055 let bits = if ty.is_ptr() { ADDRESS_BITS } else { ty.bits() };
3056 if (!ty.is_int() && !ty.is_ptr()) || !matches!(bits, 8 | 16 | 32 | 64) {
3057 return Err(self.unsupported(inst));
3058 }
3059
3060 let base = self.reg_of(addr)?;
3061 let want = self.reg_of(expected)?;
3062 let put = self.reg_of(desired)?;
3063 let got = self.new_reg(old);
3064 let flag = self.new_reg(exchanged);
3065
3066 let name = format!("cmpxchg_{bits}");
3067 let form = x86_64::form(&name).ok_or_else(|| self.unsupported(inst))?;
3068 let block = self.at.expect("a block is being filled");
3069 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
3070 let (span, flags) = (self.source.span(inst), self.carried(inst));
3071 let mut build = self.out.build(block, opcode).at(span).flags(flags);
3072 for (desc, reg) in form.operands().iter().zip([got, flag, want, put]) {
3073 let operand = mir::Operand {
3074 reg,
3075 class: desc.class,
3076 role: desc.role,
3077 constraint: desc.constraint,
3078 };
3079 build = build.operand(operand);
3080 }
3081 build.mem(mir::Mem::at(mir::Operand::read(base, self.gpr))).finish();
3082 Ok(())
3083 }
3084
3085 /// One read modify write, for the three operations this machine does in a single instruction.
3086 ///
3087 /// What the IR asks for is: read what is at an address, do something to it, put the answer back,
3088 /// say what was there before, and let nothing get between the three steps. The machine has
3089 /// `xchg` for putting a value there and `lock xadd` for adding one, and both leave what they
3090 /// found in the register the operand arrived in, which is why the value that comes back and the
3091 /// value that went in are one register here.
3092 ///
3093 /// A subtraction is the add over the negated operand, which is right at every width because the
3094 /// machine's arithmetic wraps and negating then adding is subtracting in two's complement
3095 /// whatever the operands were. The negate is a separate instruction in front, over a register of
3096 /// its own, so that the value the program handed over is not the one written on: an operand may
3097 /// be live after this and a program that read it again would read the negation.
3098 ///
3099 /// The ordering is not read, for the reason the compare and exchange beside this does not read
3100 /// it. `xchg` with memory locks the bus whether it is asked to or not and `lock xadd` is asked
3101 /// to, so both are full barriers on this machine and there is nothing weaker to fall to.
3102 ///
3103 /// Eight of the other ten never arrive, because `crate::retry` turned each of them into a loop
3104 /// around a compare and exchange before anything here saw it. The two that do arrive are the
3105 /// ones on floating values, and they are refused: a compare and exchange of a float wants the
3106 /// value carried through an integer of the same width, and an eighty bit float has no such
3107 /// width. Neither family of builtins can write one yet either, so a program that reaches this
3108 /// refusal is a program that reached an unimplemented builtin first.
3109 fn modify(&mut self, inst: Inst) -> Result<(), Unsupported> {
3110 let Extra::Rmw(op, _) = self.source[inst].extra else {
3111 return Err(self.unsupported(inst));
3112 };
3113 let args: Vec<Value> = self.source[self.source[inst].args].to_vec();
3114 let [addr, operand] = args[..] else { return Err(self.unsupported(inst)) };
3115 let old = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
3116
3117 // A value the machine can exchange in one instruction, which is an integer at one of the
3118 // four widths it has these for. A pointer arrives as an address, so it is an integer by the
3119 // time it is here, and anything else is a type this has no instruction for.
3120 let ty = self.source[old].ty;
3121 if !ty.is_int() || !matches!(ty.bits(), 8 | 16 | 32 | 64) {
3122 return Err(self.unsupported(inst));
3123 }
3124 let name = match op {
3125 RmwOp::Xchg => format!("xchg_{}", ty.bits()),
3126 RmwOp::Add | RmwOp::Sub => format!("xadd_{}", ty.bits()),
3127 _ => return Err(self.unsupported(inst)),
3128 };
3129
3130 let base = self.reg_of(addr)?;
3131 let mut put = self.reg_of(operand)?;
3132 let block = self.at.expect("a block is being filled");
3133 let span = self.source.span(inst);
3134 if op == RmwOp::Sub {
3135 let negated = self.out.new_vreg(self.gpr);
3136 let negate =
3137 mir::Opcode::new(self.names.intern(&format!("{PREFIX}neg_r_{}", ty.bits())));
3138 let form = x86_64::form(&format!("neg_r_{}", ty.bits()))
3139 .ok_or_else(|| self.unsupported(inst))?;
3140 let mut build = self.out.build(block, negate).at(span);
3141 for (desc, reg) in form.operands().iter().zip([negated, put]) {
3142 build = build.operand(mir::Operand {
3143 reg,
3144 class: desc.class,
3145 role: desc.role,
3146 constraint: desc.constraint,
3147 });
3148 }
3149 build.finish();
3150 put = negated;
3151 }
3152
3153 let got = self.new_reg(old);
3154 let form = x86_64::form(&name).ok_or_else(|| self.unsupported(inst))?;
3155 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
3156 let flags = self.carried(inst);
3157 let mut build = self.out.build(block, opcode).at(span).flags(flags);
3158 for (desc, reg) in form.operands().iter().zip([got, put]) {
3159 build = build.operand(mir::Operand {
3160 reg,
3161 class: desc.class,
3162 role: desc.role,
3163 constraint: desc.constraint,
3164 });
3165 }
3166 build.mem(mir::Mem::at(mir::Operand::read(base, self.gpr))).finish();
3167 Ok(())
3168 }
3169
3170 /// One `asm` statement.
3171 ///
3172 /// An empty template is most of the inline assembly in a test suite, and it is not a corner
3173 /// case somebody wrote by accident. A program that wants a value computed where it stands, or a
3174 /// loop the optimizer must not touch, writes `asm volatile ("" : : : "memory")`, and forty
3175 /// years of bug reports about optimizers are full of them. What such a statement asks for is
3176 /// the barrier and the operand places, and no instructions at all.
3177 ///
3178 /// So the operands are the half that is always real: a constraint says where a value has to be,
3179 /// and where it has to be is still true when the template between them is empty.
3180 ///
3181 /// What the constraints ask for, on an empty template, is only ever that two operands share a
3182 /// place. Nothing reads a register no text names, so `"r"` on its own asks for a register and
3183 /// no particular one, and any register at all answers it. A matching constraint is different,
3184 /// because it says the output the assembly leaves is the place the input arrived in, and with
3185 /// no instructions between them that is the input unchanged. So it is a rename and not a move:
3186 /// the value is already in a register and the result is that register.
3187 ///
3188 /// An output nothing is tied to and no instruction writes is whatever the assembly left there,
3189 /// which for a template that writes nothing is whatever was in the register. That is a value
3190 /// the program is not entitled to, and this writes a zero rather than reading one, because the
3191 /// allocator has to be given a definition before a use whatever the program is entitled to.
3192 ///
3193 /// # A template with instructions in it
3194 ///
3195 /// [`x86_64::read`] turns the text into the opcodes this backend already has, which is what
3196 /// `spec/11-asm-objects-debug.md` section 11.1 asks for: the machine is described once, and an
3197 /// instruction a program wrote is looked up in that description rather than copied through to
3198 /// an assembler that has one of its own. So nothing here assembles anything. What it does is
3199 /// put the statement's operands where the opcode holds them, and from there an `asm` statement
3200 /// is ordinary machine code: the allocator picks the registers, the listing and the object file
3201 /// are written from the same table as every other instruction, and a spill around one works
3202 /// because there is nothing left about it for a spill to get wrong.
3203 ///
3204 /// A register the template named in its own text is the one thing in there that is nobody's
3205 /// operand, and it is placed as itself. See [`Self::itself`] for why that is safer here than
3206 /// the thing gcc does, which is to copy the name out and leave the allocator none the wiser.
3207 ///
3208 /// Two things are refused, both for one reason, which is that placing them by a guess gives a
3209 /// program that assembles into something other than what it says.
3210 ///
3211 /// An output the template writes more than once, which is one place with two definitions in it,
3212 /// and the machine IR between here and the allocator has one definition per register by
3213 /// construction. An output tied to an input and written once is not that: it is two registers
3214 /// the description ties together, which is what [`Place`] is about.
3215 ///
3216 /// An operand read where the opcode writes, or written where it reads. An output that has not
3217 /// been written yet is not a value, and an input the assembly writes over is a value something
3218 /// else may still be using.
3219 ///
3220 /// # A register the instruction uses without being told
3221 ///
3222 /// An instruction may reach a register its text does not name, and `cpuid` is all of them at
3223 /// once: the leaf goes in `eax`, the subleaf in `ecx`, and the answer comes back in all four
3224 /// registers. The description holds every bit of that already, so what is left is to say which
3225 /// of the statement's operands is in each of those registers, and the constraint letter is the
3226 /// one thing in an assembly statement that says it. `"=a"` is an output in `rax` and `"c"` is
3227 /// an input in `rcx`, which is why a program writing `cpuid` writes its constraints that way
3228 /// and has no choice about it.
3229 ///
3230 /// A register no letter named is one the statement put nothing in, and that is the usual case
3231 /// rather than an unusual one, since an instruction that answers four questions is written by
3232 /// programs that asked one. A write of one is the register being destroyed and gets a register
3233 /// of its own, which is what tells the allocator to keep everything else out of it. A read of
3234 /// one is a register the instruction looks at and the program never filled, which gets a zero
3235 /// for the reason [`Self::undefined`] gives.
3236 ///
3237 /// # The clobber list
3238 ///
3239 /// Read now, as the registers it names being written by every instruction of the template. By
3240 /// every one rather than by one of them, because the list says the assembly as a whole leaves
3241 /// them ruined and nothing here knows which line did it. Every entry has to be a register this
3242 /// machine has a name for or the statement is refused, since a name nobody read is a register
3243 /// nobody is keeping out of.
3244 ///
3245 /// `memory` and `cc` are the two entries that are not registers and both are skipped. `memory`
3246 /// says the assembly touches storage, which is already true of every `asm` this writes and is
3247 /// nothing a register list could hold. `cc` says it ruins the condition flags, and the flag
3248 /// tracking already has that from the instructions the template was read into, since it takes
3249 /// every instruction it does not recognize as writing them and every instruction here is one
3250 /// this machine describes. `flags` is the name gcc's own register table gives the same thing on
3251 /// this machine, so a program writing it has written `cc` and is read that way: tcc's
3252 /// `tests/tcctest.c` lists both on one statement.
3253 ///
3254 /// A clobber the instruction already writes is left off it. `cpuid` writes all four registers
3255 /// by description, and a statement listing three of them as clobbers as well is saying the
3256 /// same thing twice, which the allocator would read as one register with two definitions.
3257 ///
3258 /// On a template with nothing in it the list is ignored, as it was before, since a template
3259 /// with no instructions ruins nothing whatever it said about what it ruins.
3260 fn assembly(&mut self, inst: Inst) -> Result<(), Unsupported> {
3261 let data = &self.source[inst];
3262 let Extra::Asm(asm) = data.extra else { return Err(self.unsupported(inst)) };
3263 let info = self.source[asm];
3264 if !self.source[info.targets].is_empty() {
3265 return Err(Unsupported::Assembly { inst, refused: Written::Goto });
3266 }
3267 let refused = || Unsupported::Assembly { inst, refused: Written::Operand };
3268
3269 let constraints = self.names.resolve(info.constraints).to_string();
3270 let results: Vec<Value> = data.results().collect();
3271 let operands = AsmOperands::read(&constraints, &results, &self.source[data.args])
3272 .ok_or_else(refused)?;
3273 let list: Vec<AsmOperand<'_>> = operands.iter().copied().collect();
3274
3275 // Read after the constraints and not before them, because a mnemonic whose suffix the
3276 // program left off is read at the width of the operands it names, and the operands are
3277 // what the constraints are a list of.
3278 let widths: Vec<Option<x86_64::Width>> = list
3279 .iter()
3280 .map(|operand| {
3281 let ty = self.source[operand.result.or(operand.value)?].ty;
3282 if !ty.is_scalar() {
3283 return None;
3284 }
3285 x86_64::Width::of_bits(held_bits(ty))
3286 })
3287 .collect();
3288 // An operand in memory is an address the statement holds and an object the template names,
3289 // so the reader is told which ones those are and spells `%0` for one as the object.
3290 let memory: Vec<bool> = list.iter().map(|operand| operand.memory).collect();
3291 let template = self.names.resolve(info.template).to_string();
3292 let steps = if template.trim().is_empty() {
3293 Vec::new()
3294 } else {
3295 match x86_64::read_in(&template, &widths, &memory) {
3296 Some(steps) => steps,
3297 None => return self.kept(inst, &template, &list),
3298 }
3299 };
3300
3301 // Which operands the template writes, counted before anything is placed, because the answer
3302 // decides where each of the three below comes from and one instruction may name an operand
3303 // that a later one writes. Which of them any instruction puts in a register at all is
3304 // counted in the same walk, since an operand no instruction reaches that way is one nothing
3305 // has to put anywhere: a constant a template names only as the distance into an address is
3306 // written into the instruction, and a register holding a copy of it would be one nobody
3307 // reads. An operand the address is counted from is reached that way and is counted here for
3308 // that reason, because the walk below it is over the opcode's operands and an address is
3309 // not one of those.
3310 //
3311 // Whether any instruction reads an operand an instruction above it wrote is counted in the
3312 // same walk too. Such a template is one whose instructions have to be written in order with
3313 // each read taken from wherever the last write left the operand, which is what
3314 // [`Self::woven`] does, and so is one that writes an operand twice.
3315 let mut writes = vec![0usize; list.len()];
3316 let mut reads = vec![false; list.len()];
3317 let mut held = vec![false; list.len()];
3318 let mut after = false;
3319 for step in &steps {
3320 // A call out of the template writes every register the convention lets the callee
3321 // leave anything in, and an output pinned to one of those is written by it.
3322 if let x86_64::Step::Call { .. } = step {
3323 for index in self.lost(&list).into_iter().filter_map(|(_, _, index)| index) {
3324 *writes.get_mut(index).ok_or_else(refused)? += 1;
3325 }
3326 continue;
3327 }
3328 let x86_64::Step::Line(line) = step else { continue };
3329 match line.at.and_then(|at| at.base) {
3330 Some(x86_64::Piece::Operand { index, .. }) => {
3331 *held.get_mut(index).ok_or_else(refused)? = true;
3332 after |= writes[index] > 0;
3333 }
3334 Some(x86_64::Piece::Reg { reg, .. }) => {
3335 if let Some(index) = bound(&list, reg, Role::Use) {
3336 *held.get_mut(index).ok_or_else(refused)? = true;
3337 after |= writes[index] > 0;
3338 }
3339 }
3340 _ => {}
3341 }
3342 let mut written = Vec::new();
3343 let form = x86_64::form(line.opcode).ok_or_else(refused)?;
3344 // Which registers the instruction reaches, asked the same way it is asked again when
3345 // the instruction is written. See [`Self::lettered`] for the one opcode whose answer
3346 // comes from the constraint letters rather than from the description.
3347 let lettered = (line.opcode == x86_64::LITERAL).then(|| self.lettered(&list));
3348 let (described, pieces) = match &lettered {
3349 Some((described, pieces)) => (described.as_slice(), pieces.as_slice()),
3350 None => (form.operands(), line.operands.as_slice()),
3351 };
3352 for (desc, piece) in described.iter().zip(pieces) {
3353 // An operand the instruction reaches without its text saying so is the statement's
3354 // only when a constraint letter put something there. One that is nobody's writes
3355 // nothing of the program's, so it is counted nowhere and is dealt with where it is
3356 // placed.
3357 let index = match *piece {
3358 x86_64::Piece::Operand { index, .. } => index,
3359 x86_64::Piece::Implicit { reg } => match bound(&list, reg, desc.role) {
3360 Some(index) => index,
3361 None => continue,
3362 },
3363 x86_64::Piece::Reg { reg, .. } => match bound(&list, reg, desc.role) {
3364 Some(index) => index,
3365 None => continue,
3366 },
3367 };
3368 *held.get_mut(index).ok_or_else(refused)? = true;
3369 if matches!(desc.role, Role::Def | Role::EarlyDef) {
3370 written.push(index);
3371 } else {
3372 *reads.get_mut(index).ok_or_else(refused)? = true;
3373 after |= writes[index] > 0;
3374 }
3375 }
3376 for index in written {
3377 *writes.get_mut(index).ok_or_else(refused)? += 1;
3378 }
3379 }
3380 let woven = after
3381 || writes.iter().any(|&count| count > 1)
3382 || steps.iter().any(|step| !matches!(step, x86_64::Step::Line(_)));
3383
3384 // Where every operand is. Worked out in full before the first instruction is written, since
3385 // reading a value may be what puts it in a register in the first place, and that has to
3386 // happen in front of the assembly rather than in the middle of it.
3387 let mut places: Vec<Place> = vec![Place::default(); list.len()];
3388 for (index, operand) in list.iter().copied().enumerate() {
3389 let Some(result) = operand.result else {
3390 // An input, or an output the assembly was handed the address of, and both are a
3391 // value that arrives in a register and is read out of it, unless no instruction of
3392 // the template reads it out of one.
3393 let value = operand.value.ok_or_else(refused)?;
3394 if held[index] {
3395 places[index].read = Some(self.reg_of(value)?);
3396 }
3397 continue;
3398 };
3399 let ty = self.source[result].ty;
3400 if on_x87(ty) {
3401 return Err(refused());
3402 }
3403 let tied = operands.tied_to(index);
3404 if let Some(from) = tied {
3405 if self.class_of(self.source[from].ty) != self.class_of(ty) {
3406 return Err(refused());
3407 }
3408 places[index].read = Some(self.reg_of(from)?);
3409 }
3410 if writes[index] > 0 {
3411 places[index].write = Some(self.new_reg(result));
3412 continue;
3413 }
3414 match tied {
3415 // The place the input arrived in, which the assembly wrote nothing over. One
3416 // register, so this is a rename rather than a move.
3417 Some(_) => {
3418 let reg = places[index].read.ok_or_else(refused)?;
3419 self.regs[result.index()] = Some(reg);
3420 places[index].write = Some(reg);
3421 }
3422 None => {
3423 self.undefined(inst, result)?;
3424 places[index].write = self.regs[result.index()];
3425 }
3426 }
3427 }
3428
3429 // An output an instruction of the template also reads, which the statement said nothing
3430 // about because an output is what a statement says the other thing about. What it holds
3431 // there is undefined, and a program writing one means it: `sbb %0, %0` in libgmp's
3432 // `add_mssaaaa` subtracts a register from itself and is asking for the borrow bit rather
3433 // than for the number, so whatever the register held, the answer is the same. Undefined is
3434 // not the same as absent though, since the allocator is owed a definition in front of every
3435 // use, so it gets the zero an output nothing wrote gets and for the same reason.
3436 //
3437 // Unless an input could have been in the same register, in which case gcc's allocator puts
3438 // it there whenever it can and a program may have been written against that. tcc's test of
3439 // a call from a template reads its output `"=a" (s)` to pass `"r" (str)` to `getenv`, which
3440 // is only the string because gcc gave the two of them `rax`. So an output nothing has
3441 // written yet reads the one input that could share its place, when there is exactly one.
3442 // One written `&` is written before the inputs are read and shares nothing.
3443 for index in 0..list.len() {
3444 if !reads[index] || places[index].read.is_some() || places[index].write.is_none() {
3445 continue;
3446 }
3447 let reg = match self.shared(&list, index) {
3448 Some(value) => self.reg_of(value)?,
3449 None => self.seeded(inst, list[index])?,
3450 };
3451 places[index].read = Some(reg);
3452 }
3453
3454 // Worked out once for the whole template, since the list is one list and every instruction
3455 // of the template gets it. Not worked out at all for a template with no instructions, which
3456 // is where there is nothing for it to go on.
3457 let clobbers = self.names.resolve(info.clobbers).to_string();
3458 let clobbered =
3459 if steps.is_empty() { Vec::new() } else { Self::clobbered(inst, &clobbers)? };
3460
3461 // A template with a label in it is not one run of instructions, and what it is instead is
3462 // in [`Self::woven`], which is also where a template goes whose instructions read what the
3463 // ones above them wrote. Every other template is what it has always been, which is every
3464 // instruction of it written into the block the statement stands in.
3465 if woven {
3466 return self.woven(inst, &steps, &mut places, &list, &clobbered, &writes);
3467 }
3468 for step in &steps {
3469 let x86_64::Step::Line(line) = step else { continue };
3470 self.instruction(inst, line, &places, &list, &clobbered)?;
3471 }
3472 Ok(())
3473 }
3474
3475 /// A template the reader could not take apart, kept as its text. See [`x86_64::Form::Template`].
3476 ///
3477 /// What the text names is spelled into it here, the way gcc prints it into its listing: a
3478 /// constant as `$5`, or as `5` under the `c` modifier, and the address of a name as the name.
3479 /// An object in memory is the one thing that cannot be spelled yet, since where it is depends on
3480 /// registers nothing has chosen, so it is left as a hole the writer fills and its address is the
3481 /// instruction's memory operand. One is all an instruction has room for, and every template this
3482 /// has met names one at most. An operand in a register is refused for now, as is a template
3483 /// that names one by name rather than by number.
3484 ///
3485 /// A statement written with no colons is basic assembly, where `%` is a character like any
3486 /// other and a register is written `%eax`. The front end keeps no mark of which kind a statement
3487 /// was, so one with no operands and no clobbers is read as basic, which is what gcc would do for
3488 /// every such template but one written with empty colons around it.
3489 ///
3490 /// The registers a call may write are taken as written, see below for why.
3491 fn kept(
3492 &mut self,
3493 inst: Inst,
3494 template: &str,
3495 list: &[AsmOperand<'_>],
3496 ) -> Result<(), Unsupported> {
3497 // Refused as the template it is, since keeping it is what was tried after reading it
3498 // failed, and what could not be kept is what it names rather than any one operand.
3499 let refused = || Unsupported::Assembly { inst, refused: Written::Template };
3500 let data = &self.source[inst];
3501 let Extra::Asm(asm) = data.extra else { return Err(self.unsupported(inst)) };
3502 let clobbers = self.names.resolve(self.source[asm].clobbers).to_string();
3503 let basic = list.is_empty() && clobbers.trim().is_empty();
3504
3505 let mut text = String::with_capacity(template.len());
3506 let mut memory: Option<usize> = None;
3507 if basic {
3508 text.push_str(template);
3509 } else {
3510 let mut chars = template.chars().peekable();
3511 // Inside `{att|intel}`, and past the `|` in it, which is the half nobody reads.
3512 let mut dialect = false;
3513 let mut skipped = false;
3514 while let Some(c) = chars.next() {
3515 match c {
3516 '{' => {
3517 dialect = true;
3518 continue;
3519 }
3520 '|' if dialect => {
3521 skipped = true;
3522 continue;
3523 }
3524 '}' if dialect => {
3525 dialect = false;
3526 skipped = false;
3527 continue;
3528 }
3529 _ if skipped => continue,
3530 '%' => {}
3531 _ => {
3532 text.push(c);
3533 continue;
3534 }
3535 }
3536 match chars.peek().copied() {
3537 Some(c @ ('%' | '{' | '|' | '}')) => {
3538 chars.next();
3539 text.push(c);
3540 continue;
3541 }
3542 Some('=') => {
3543 chars.next();
3544 text.push_str(&inst.index().to_string());
3545 continue;
3546 }
3547 _ => {}
3548 }
3549 let modifier = match chars.peek().copied() {
3550 Some(c) if c.is_ascii_alphabetic() => {
3551 chars.next();
3552 Some(c)
3553 }
3554 _ => None,
3555 };
3556 let mut digits = String::new();
3557 while let Some(c) = chars.peek().copied().filter(char::is_ascii_digit) {
3558 digits.push(c);
3559 chars.next();
3560 }
3561 let index: usize = digits.parse().map_err(|_| refused())?;
3562 let operand = list.get(index).ok_or_else(refused)?;
3563 if operand.memory {
3564 if modifier.is_some() || memory.is_some_and(|had| had != index) {
3565 return Err(refused());
3566 }
3567 memory = Some(index);
3568 text.push_str(x86_64::TEMPLATE_MEM);
3569 continue;
3570 }
3571 if operand.result.is_some() {
3572 return Err(refused());
3573 }
3574 let value = operand.value.ok_or_else(refused)?;
3575 let bare = match modifier {
3576 None => false,
3577 Some('c' | 'P' | 'p') => true,
3578 Some(_) => return Err(refused()),
3579 };
3580 if !bare {
3581 text.push('$');
3582 }
3583 if let Some(number) = self.number(value) {
3584 text.push_str(&number.to_string());
3585 } else if let Some(symbol) = self.named_address(value) {
3586 text.push_str(&x86_64::template_name(self.names.resolve(symbol)));
3587 } else {
3588 return Err(refused());
3589 }
3590 }
3591 }
3592
3593 // Every register a call may leave anything in, as well as the ones the list names. The
3594 // text can write any register it likes without saying so, and tcc's tests do: gcc gets
3595 // away with that at `-O0` because nothing lives in a register between two statements
3596 // there, and taking these away from the allocator across the template is what gives the
3597 // same answer here. Nothing is written to them by this, so a register one template leaves
3598 // a value in is still holding it when the next template reads it.
3599 let mut clobbered: Vec<(PhysReg, RegClass)> =
3600 self.lost(list).into_iter().map(|(reg, class, _)| (reg, class)).collect();
3601 for reg in Self::clobbered(inst, &clobbers)? {
3602 if !clobbered.iter().any(|&(had, _)| had == reg) {
3603 clobbered.push((reg, self.gpr));
3604 }
3605 }
3606 // An object in this function's frame is named by where it is in the frame, the way gcc
3607 // names it, rather than by a register its address was put in first. The text may write
3608 // registers it does not declare, and tcc's tests do: one that writes `%ecx` behind the
3609 // compiler's back would otherwise take the address with it.
3610 let mut local = None;
3611 let at = match memory {
3612 Some(index) => {
3613 let value = list[index].value.ok_or_else(refused)?;
3614 local = self.local_of(value);
3615 let base = match local {
3616 Some(_) => mir::Reg::physical(self.conv.stack_pointer),
3617 None => self.reg_of(value)?,
3618 };
3619 Some(mir::Mem::at(mir::Operand::read(base, self.gpr)))
3620 }
3621 None => None,
3622 };
3623 let symbol = self.names.intern(&text);
3624 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::TEMPLATE)));
3625 let block = self.at.expect("a block is being filled");
3626 let span = self.source.span(inst);
3627 let mut build = self.out.build(block, opcode).at(span).symbol(symbol);
3628 for (reg, class) in clobbered {
3629 build = build.operand(mir::Operand::write(mir::Reg::physical(reg), class));
3630 }
3631 if let Some(mem) = at {
3632 build = build.mem(mem);
3633 }
3634 let made = build.finish();
3635 if let Some(local) = local {
3636 self.stack.addresses.push((made, local));
3637 }
3638 Ok(())
3639 }
3640
3641 /// The object in this function's frame a value is the address of, for one an `alloca` of a
3642 /// size known here made. See [`Self::reserve`], which is where the `lea` it is found by came
3643 /// from.
3644 fn local_of(&self, value: Value) -> Option<usize> {
3645 let Def::Result { inst, .. } = self.source[value].def else { return None };
3646 if self.source[inst].opcode != Opcode::Alloca
3647 || !self.source[self.source[inst].args].is_empty()
3648 {
3649 return None;
3650 }
3651 let reg = self.regs[value.index()]?;
3652 self.stack.addresses.iter().find_map(|&(made, local)| {
3653 let data = &self.out[made];
3654 let defined = self.out[data.operands].first()?;
3655 (defined.reg == reg).then_some(local)
3656 })
3657 }
3658
3659 /// The name a value is the address of, for one a `global_addr` defined.
3660 fn named_address(&self, value: Value) -> Option<Symbol> {
3661 let Def::Result { inst, .. } = self.source[value].def else { return None };
3662 if self.source[inst].opcode != Opcode::GlobalAddr {
3663 return None;
3664 }
3665 let Extra::Symbol(symbol) = self.source[inst].extra else { return None };
3666 Some(symbol)
3667 }
3668
3669 /// A register holding a zero, for an operand of a template that is read before anything filled
3670 /// it.
3671 ///
3672 /// Two things ask for this and they are the same thing twice. An output the template reads has
3673 /// nothing to be read out of until the instruction that writes it has run, and a loop carries
3674 /// an operand into a block before the instruction that fills it, so both are a use in front of
3675 /// every definition. What the program is owed there is nothing, since the value is undefined
3676 /// either way, and what the allocator is owed is a register something wrote.
3677 fn seeded(&mut self, inst: Inst, operand: AsmOperand<'_>) -> Result<mir::Reg, Unsupported> {
3678 let refused = || Unsupported::Assembly { inst, refused: Written::Operand };
3679 let value = operand.result.or(operand.value).ok_or_else(refused)?;
3680 let class = self.class_of(self.source[value].ty);
3681 if class != self.gpr {
3682 return Err(refused());
3683 }
3684 let block = self.at.expect("a block is being filled");
3685 let reg = self.out.new_vreg(class);
3686 let put = mir::Opcode::new(self.names.intern(&format!("{PREFIX}mov_ri_64")));
3687 self.out.build(block, put).at(self.source.span(inst)).def(reg, class).imm(0).finish();
3688 Ok(reg)
3689 }
3690
3691 /// A template with labels in it, as the blocks its jumps leave and arrive at.
3692 ///
3693 /// A statement is an instruction of the IR and stands inside one block, so a template that
3694 /// jumps has to stop being one thing. Each label becomes a block, each jump ends the block it
3695 /// stands in and gives it two arms, and whatever follows the statement goes into whichever
3696 /// block the walk finished in, which is what [`Self::block`] already reads off `self.at` and
3697 /// what [`Self::saves_place`] already does for the same reason.
3698 ///
3699 /// # What is carried between them
3700 ///
3701 /// The machine IR here is in the form where a register is written once, so an operand written
3702 /// inside a loop and read again at the top of it cannot be one register. What arrives at the
3703 /// top is a parameter of that block, and every jump to it carries whichever register held the
3704 /// operand where the jump stands. That is the whole of the bookkeeping: every block a label
3705 /// made takes one parameter for each operand that is in a register at all, in one order, so an
3706 /// arm's arguments and a block's parameters are the same list read twice.
3707 ///
3708 /// Which register an operand is in at each point is kept in the read half of its place, since
3709 /// that is what the instructions below read it out of. An instruction that writes an operand
3710 /// leaves it in the register it wrote, and a jump below carries that one. The block an
3711 /// untaken jump falls into is arrived at one way only and so takes no parameters, and nothing
3712 /// about where the operands are changes there.
3713 ///
3714 /// An operand written by the template and filled by nothing is written as a zero first, for
3715 /// the reason [`Self::undefined`] gives and one more: a jump may carry it before the
3716 /// instruction that fills it has run, and an argument has to be a register something wrote.
3717 ///
3718 /// # The condition state
3719 ///
3720 /// Nothing carries it and nothing has to. The instruction that sets it and the jump that reads
3721 /// it are both written here, next to each other in one block, and what the allocator may put
3722 /// between them is a move, which on this machine leaves the condition state alone. The edge
3723 /// into a block a loop goes back to is a critical edge and `crate::split` gives it a block of
3724 /// its own, so the moves an arm turns into land behind the jump rather than in front of it.
3725 fn woven(
3726 &mut self,
3727 inst: Inst,
3728 steps: &[x86_64::Step],
3729 places: &mut [Place],
3730 list: &[AsmOperand<'_>],
3731 clobbered: &[PhysReg],
3732 writes: &[usize],
3733 ) -> Result<(), Unsupported> {
3734 let refused = || Unsupported::Assembly { inst, refused: Written::Operand };
3735 let span = self.source.span(inst);
3736
3737 // Which operands are carried, which is every one that is in a register at all. An operand
3738 // the template never puts in one, such as a constant it names only as the distance into an
3739 // address, is in the instruction and has nowhere to be carried from.
3740 let mut carried: Vec<(usize, RegClass)> = Vec::new();
3741 for (index, operand) in list.iter().enumerate() {
3742 if places[index].read.is_none() && places[index].write.is_none() {
3743 continue;
3744 }
3745 let value = operand.result.or(operand.value).ok_or_else(refused)?;
3746 let ty = self.source[value].ty;
3747 if on_x87(ty) {
3748 return Err(refused());
3749 }
3750 carried.push((index, self.class_of(ty)));
3751 }
3752
3753 // What each of them holds where the template starts.
3754 for &(index, _) in &carried {
3755 if places[index].read.is_some() {
3756 continue;
3757 }
3758 if writes[index] == 0 {
3759 places[index].read = places[index].write;
3760 continue;
3761 }
3762 places[index].read = Some(self.seeded(inst, list[index])?);
3763 }
3764
3765 // The blocks, made before the walk because a jump forwards names a label the walk has not
3766 // reached yet.
3767 let mut labels: Vec<(&str, mir::Block, Vec<mir::Reg>)> = Vec::new();
3768 for step in steps {
3769 let x86_64::Step::Label(name) = step else { continue };
3770 let block = self.out.create_block();
3771 let mut params = Vec::with_capacity(carried.len());
3772 for &(_, class) in &carried {
3773 params.push(self.out.append_param(block, class));
3774 }
3775 labels.push((name.as_str(), block, params));
3776 }
3777
3778 let mut wrote: Vec<usize> = Vec::new();
3779 for step in steps {
3780 match step {
3781 x86_64::Step::Label(name) => {
3782 let (block, params) = Self::went(&labels, name).ok_or_else(refused)?;
3783 let from = self.at.expect("a block is being filled");
3784 let args = Self::held(places, &carried).ok_or_else(refused)?;
3785 *self.out.succs_mut(from) = vec![mir::BlockCall::with(block, args)];
3786 self.at = Some(block);
3787 for (at, &(index, _)) in carried.iter().enumerate() {
3788 places[index].read = params.get(at).copied();
3789 }
3790 }
3791 x86_64::Step::Jump { opcode, to } => {
3792 let (block, _) = Self::went(&labels, to).ok_or_else(refused)?;
3793 let from = self.at.expect("a block is being filled");
3794 let args = Self::held(places, &carried).ok_or_else(refused)?;
3795 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{opcode}")));
3796 self.out.build(from, opcode).at(span).finish();
3797 let next = self.out.create_block();
3798 *self.out.succs_mut(from) =
3799 vec![mir::BlockCall::with(block, args), mir::BlockCall::to(next)];
3800 self.at = Some(next);
3801 }
3802 x86_64::Step::Away { symbol } => {
3803 // Only in a function that is written without a prologue, which is the one
3804 // place the jump means what it says. Anywhere else there is an epilogue behind
3805 // the statement that puts the registers back and gives the frame up, and a
3806 // jump over it goes to the next function with this function's frame still
3807 // taken. The reader already made sure it is the last step of the template, so
3808 // what is left to ask is about the function around it.
3809 if !self.source.attrs.set.contains(AttrSet::NAKED) {
3810 return Err(Unsupported::Assembly { inst, refused: Written::Away });
3811 }
3812 let from = self.at.expect("a block is being filled");
3813 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{AWAY}")));
3814 let symbol = self.names.intern(symbol);
3815 self.out.build(from, opcode).at(span).symbol(symbol).finish();
3816 // Nowhere, which is what a jump out of the function leaves behind it and is
3817 // the same list a `ret` leaves. The block after it is made for the walk above
3818 // rather than for the program: the statement may be in the middle of a body
3819 // that goes on being lowered, and what that lowering writes is reached by
3820 // nothing and thrown away with the block.
3821 *self.out.succs_mut(from) = Vec::new();
3822 self.at = Some(self.out.create_block());
3823 }
3824 x86_64::Step::Call { symbol } => {
3825 self.call_out(inst, symbol, places, list, clobbered, &carried, &mut wrote)?;
3826 }
3827 x86_64::Step::Line(line) => {
3828 let form = x86_64::form(line.opcode).ok_or_else(refused)?;
3829 let mut written = Vec::new();
3830 for (desc, piece) in form.operands().iter().zip(&line.operands) {
3831 if !desc.role.is_def() {
3832 continue;
3833 }
3834 let index = match *piece {
3835 x86_64::Piece::Operand { index, .. } => index,
3836 x86_64::Piece::Implicit { reg } => match bound(list, reg, desc.role) {
3837 Some(index) => index,
3838 None => continue,
3839 },
3840 x86_64::Piece::Reg { reg, .. } => match bound(list, reg, desc.role) {
3841 Some(index) => index,
3842 None => continue,
3843 },
3844 };
3845 written.push(index);
3846 }
3847 // A register is written once in this form of the machine IR, so an operand
3848 // an instruction above already wrote is written into a new one here, and what
3849 // reads it below reads that one.
3850 for &index in &written {
3851 if !wrote.contains(&index) {
3852 wrote.push(index);
3853 continue;
3854 }
3855 let &(_, class) =
3856 carried.iter().find(|&&(at, _)| at == index).ok_or_else(refused)?;
3857 let place = places.get_mut(index).ok_or_else(refused)?;
3858 place.write = Some(self.out.new_vreg(class));
3859 }
3860 self.instruction(inst, line, places, list, clobbered)?;
3861 for index in written {
3862 let place = places.get_mut(index).ok_or_else(refused)?;
3863 if place.write.is_some() {
3864 place.read = place.write;
3865 }
3866 }
3867 }
3868 }
3869 }
3870
3871 // Where the walk left each output, which is the parameter of the block a label made when
3872 // the template ends in one and the register an instruction wrote when it does not.
3873 for (index, operand) in list.iter().enumerate() {
3874 let Some(result) = operand.result else { continue };
3875 if let Some(reg) = places[index].read {
3876 self.regs[result.index()] = Some(reg);
3877 }
3878 }
3879 Ok(())
3880 }
3881
3882 /// A template's call to a function somewhere else, as the call the convention makes.
3883 ///
3884 /// The opcode is the one a call written in C becomes, so everything that asks whether a
3885 /// function calls anything gets the answer it would for one: the stack pointer is left aligned
3886 /// at the statement and nothing is kept in the red zone. What is not the same is the operands.
3887 /// Nothing is passed by the convention, since the template put the arguments where it wanted
3888 /// them, and what comes back is whatever an output is pinned to, since that is the only thing
3889 /// the template says about it. Every other register the callee may leave anything in is
3890 /// written here, which is what a program that calls from a template never says and always
3891 /// means.
3892 #[allow(clippy::too_many_arguments)]
3893 fn call_out(
3894 &mut self,
3895 inst: Inst,
3896 symbol: &str,
3897 places: &mut [Place],
3898 list: &[AsmOperand<'_>],
3899 clobbered: &[PhysReg],
3900 carried: &[(usize, RegClass)],
3901 wrote: &mut Vec<usize>,
3902 ) -> Result<(), Unsupported> {
3903 let refused = || Unsupported::Assembly { inst, refused: Written::Operand };
3904 let mut operands = Vec::new();
3905 let mut written = Vec::new();
3906 let lost = self.lost(list);
3907 for &(reg, class, index) in &lost {
3908 let Some(index) = index else {
3909 operands.push(mir::Operand::write(mir::Reg::physical(reg), class));
3910 continue;
3911 };
3912 // Written once in this form of the machine IR, so a second write is a new register,
3913 // the same as for an instruction in [`Self::woven`].
3914 if wrote.contains(&index) {
3915 let &(_, class) =
3916 carried.iter().find(|&&(at, _)| at == index).ok_or_else(refused)?;
3917 places.get_mut(index).ok_or_else(refused)?.write = Some(self.out.new_vreg(class));
3918 } else {
3919 wrote.push(index);
3920 }
3921 let place = places.get(index).ok_or_else(refused)?.write.ok_or_else(refused)?;
3922 operands.push(mir::Operand::write(place, class).with(Constraint::Fixed(reg)));
3923 written.push(index);
3924 }
3925 for ® in clobbered {
3926 if lost.iter().all(|&(gone, class, _)| gone != reg || class != self.gpr) {
3927 operands.push(mir::Operand::write(mir::Reg::physical(reg), self.gpr));
3928 }
3929 }
3930 let block = self.at.expect("a block is being filled");
3931 let span = self.source.span(inst);
3932 let opcode = mir::Opcode::new(self.names.intern(abi::CALL));
3933 let symbol = self.names.intern(symbol);
3934 let mut build = self.out.build(block, opcode).at(span).symbol(symbol);
3935 for operand in operands {
3936 build = build.operand(operand);
3937 }
3938 build.finish();
3939 let calls = &mut self.stack.calls;
3940 *calls = Some(calls.unwrap_or(0));
3941 for index in written {
3942 let place = places.get_mut(index).ok_or_else(refused)?;
3943 place.read = place.write;
3944 }
3945 Ok(())
3946 }
3947
3948 /// Every register a call may leave anything in, with its file and the output pinned to it if
3949 /// one is.
3950 ///
3951 /// Only a general purpose register is ever pinned to an output, since those are the only ones a
3952 /// constraint letter or a register variable names here. The vector registers are numbered from
3953 /// nought as well, so asking about one of them would find the output pinned to the register of
3954 /// the same number in the other file.
3955 fn lost(&self, list: &[AsmOperand<'_>]) -> Vec<(PhysReg, RegClass, Option<usize>)> {
3956 let conv = self.conv;
3957 let ints = conv.int_order.iter().filter(|&®| !conv.preserves_int(reg));
3958 let sses = conv.sse_order.iter().filter(|&®| !conv.preserves_sse(reg));
3959 ints.map(|®| (reg, conv.int_class, bound(list, reg, Role::Def)))
3960 .chain(sses.map(|®| (reg, conv.sse_class, None)))
3961 .collect()
3962 }
3963
3964 /// The input an output read before anything wrote it shares its register with, which is the
3965 /// one input that could be in that register, or nothing when there is none or more than one.
3966 ///
3967 /// Could be means nothing ties it elsewhere: it is in a register rather than in memory, no
3968 /// constraint pins it anywhere the output is not, and it is not tied to another output. An
3969 /// output written `&` shares nothing, since the assembly writes it before it reads the inputs.
3970 fn shared(&self, list: &[AsmOperand<'_>], index: usize) -> Option<Value> {
3971 let output = list.get(index)?;
3972 if output.early || output.tied.is_some() {
3973 return None;
3974 }
3975 let class = self.class_of(self.source[output.result?].ty);
3976 let mut fits = list.iter().filter(|operand| {
3977 operand.result.is_none()
3978 && !operand.memory
3979 && operand.tied.is_none()
3980 && operand.value.is_some_and(|value| self.class_of(self.source[value].ty) == class)
3981 && pinned(operand).is_none_or(|reg| pinned(output) == Some(reg))
3982 });
3983 let value = fits.next()?.value;
3984 if fits.next().is_some() {
3985 return None;
3986 }
3987 value
3988 }
3989
3990 /// The block one of the template's labels made, and the parameters it takes.
3991 fn went<'b>(
3992 labels: &'b [(&str, mir::Block, Vec<mir::Reg>)],
3993 name: &str,
3994 ) -> Option<(mir::Block, &'b [mir::Reg])> {
3995 labels
3996 .iter()
3997 .find(|(had, ..)| *had == name)
3998 .map(|(_, block, params)| (*block, params.as_slice()))
3999 }
4000
4001 /// The register each carried operand is in, which is what an arm to a label carries.
4002 fn held(places: &[Place], carried: &[(usize, RegClass)]) -> Option<Vec<mir::Reg>> {
4003 carried.iter().map(|&(index, _)| places.get(index)?.read).collect()
4004 }
4005
4006 /// The registers a clobber list names, in the order it named them.
4007 ///
4008 /// Nothing is dropped. A name this has no register for is refused, because the list is the
4009 /// program telling the compiler which registers it may not leave anything in, and an entry
4010 /// nobody read is a register something may still be left in. See [`Self::assembly`] for the
4011 /// two entries that are not registers and for why they are skipped rather than refused.
4012 fn clobbered(inst: Inst, clobbers: &str) -> Result<Vec<PhysReg>, Unsupported> {
4013 let refused = || Unsupported::Assembly { inst, refused: Written::Clobber };
4014 let mut named = Vec::new();
4015 for entry in clobbers.split(',') {
4016 let entry = entry.trim().trim_matches('"');
4017 // The sigil is optional in a clobber list and means nothing when it is there, unlike
4018 // in a template, where it is what tells a register from an operand.
4019 let entry = entry.strip_prefix('%').unwrap_or(entry);
4020 if entry.is_empty() || matches!(entry, "memory" | "cc" | "flags") {
4021 continue;
4022 }
4023 let (reg, _) = x86_64::gpr_named(entry).ok_or_else(refused)?;
4024 if !named.contains(®) {
4025 named.push(reg);
4026 }
4027 }
4028 Ok(named)
4029 }
4030
4031 /// One instruction of a template, as the machine instruction it was read back into.
4032 fn instruction(
4033 &mut self,
4034 inst: Inst,
4035 line: &x86_64::Line,
4036 places: &[Place],
4037 list: &[AsmOperand<'_>],
4038 clobbered: &[PhysReg],
4039 ) -> Result<(), Unsupported> {
4040 let refused = || Unsupported::Assembly { inst, refused: Written::Operand };
4041 let form = x86_64::form(line.opcode).ok_or_else(refused)?;
4042 // What the instruction reaches and what is in each of them. The description answers the
4043 // first for every opcode but one, and the pieces the template was read into answer the
4044 // second. Bytes a program wrote out itself are the one, since nothing in a number is a
4045 // register anybody could read, so the constraint letters answer both. See
4046 // [`Self::lettered`].
4047 let lettered = (line.opcode == x86_64::LITERAL).then(|| self.lettered(list));
4048 let (described, pieces) = match &lettered {
4049 Some((described, pieces)) => (described.as_slice(), pieces.as_slice()),
4050 None => (form.operands(), line.operands.as_slice()),
4051 };
4052 let mut built = Vec::with_capacity(pieces.len() + clobbered.len());
4053 for (desc, piece) in described.iter().zip(pieces) {
4054 built.push(self.placed(inst, *desc, *piece, places, list)?);
4055 }
4056 // The clobbers go in among the definitions rather than behind the reads, because an operand
4057 // vector in the machine IR is every definition and then every use and what counts them
4058 // reads that order rather than each operand's role.
4059 let defs = built.iter().take_while(|operand| operand.role.is_def()).count();
4060 let mut added = 0usize;
4061 for ® in clobbered {
4062 if described.iter().any(|desc| desc.constraint == Constraint::Fixed(reg)) {
4063 continue;
4064 }
4065 built.insert(defs, mir::Operand::write(mir::Reg::physical(reg), self.gpr));
4066 added += 1;
4067 }
4068 // A constraint tying one operand to another names it by its place in this vector, and the
4069 // clobbers were put in the middle of the vector, so everything behind them moved. The
4070 // description is written against an instruction with no clobbers in it and cannot know
4071 // that, which makes this the one place the two numberings have to be reconciled.
4072 for operand in &mut built {
4073 if let Constraint::Reuse(at) = operand.constraint {
4074 if usize::from(at) >= defs {
4075 let moved = usize::from(at) + added;
4076 operand.constraint =
4077 Constraint::Reuse(u8::try_from(moved).map_err(|_| refused())?);
4078 }
4079 }
4080 }
4081 let at = match line.at {
4082 Some(at) => Some(self.addressed(inst, at, places, list)?),
4083 None => None,
4084 };
4085
4086 let block = self.at.expect("a block is being filled");
4087 let span = self.source.span(inst);
4088 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", line.opcode)));
4089 let mut build = self.out.build(block, opcode).at(span);
4090 for operand in built {
4091 build = build.operand(operand);
4092 }
4093 if let Some(value) = line.imm {
4094 build = build.imm(value);
4095 }
4096 if let Some(mem) = at {
4097 build = build.mem(mem);
4098 }
4099 build.finish();
4100 Ok(())
4101 }
4102
4103 /// The registers a run of bytes reaches, taken from the constraint letters rather than from the
4104 /// description of an opcode.
4105 ///
4106 /// Every other instruction of a template has a description saying which registers it reaches
4107 /// without naming them, and [`Self::assembly`] matches the letters against that. Bytes a program
4108 /// wrote out itself have no such description and could not have one: what the instruction is, is
4109 /// a number, and nothing in a number is a register anything could read. So the letters are the
4110 /// whole of what is known, and they are enough, because a program writing an instruction this
4111 /// way has to say where its operands go for exactly the reason a program writing `cpuid` does.
4112 ///
4113 /// Each register named by a letter gets one entry for the write and one for the read, the same
4114 /// two `cpuid` has, and only the half the statement asked for: a register no output names is not
4115 /// written here and one no input names is not read. The writes come first because that is the
4116 /// order an operand vector in the machine IR is counted in. A register named by nothing is left
4117 /// out rather than given a spare one, which is the difference from `cpuid` and is right for the
4118 /// same reason: `cpuid` writes four registers whatever the program said, and what these bytes
4119 /// touch is known only from what the program said.
4120 fn lettered(&self, list: &[AsmOperand<'_>]) -> (Vec<OperandDesc>, Vec<x86_64::Piece>) {
4121 let mut named: Vec<PhysReg> = Vec::new();
4122 for operand in list {
4123 if let Some(reg) = pinned(operand) {
4124 if !named.contains(®) {
4125 named.push(reg);
4126 }
4127 }
4128 }
4129 let mut described = Vec::with_capacity(named.len() * 2);
4130 let mut pieces = Vec::with_capacity(named.len() * 2);
4131 for role in [Role::Def, Role::Use] {
4132 for ® in &named {
4133 if bound(list, reg, role).is_none() {
4134 continue;
4135 }
4136 let desc = if role.is_def() {
4137 OperandDesc::write(self.gpr)
4138 } else {
4139 OperandDesc::read(self.gpr)
4140 };
4141 described.push(desc.with(Constraint::Fixed(reg)));
4142 pieces.push(x86_64::Piece::Implicit { reg });
4143 }
4144 }
4145 (described, pieces)
4146 }
4147
4148 /// One operand of one instruction of a template, in the register the statement put it in.
4149 fn placed(
4150 &mut self,
4151 inst: Inst,
4152 desc: OperandDesc,
4153 piece: x86_64::Piece,
4154 places: &[Place],
4155 list: &[AsmOperand<'_>],
4156 ) -> Result<mir::Operand, Unsupported> {
4157 let refused = || Unsupported::Assembly { inst, refused: Written::Operand };
4158 // A register the instruction reaches without its text naming it belongs to whichever of the
4159 // statement's operands a constraint letter put there, and to nobody when no letter did.
4160 // There is no width to check in that case: the operand is the register the letter named and
4161 // the instruction does what it does to it, which is what a program writing `"=a"` asked for.
4162 let (index, spelled) = match piece {
4163 x86_64::Piece::Operand { index, width, stated } => (index, Some((width, stated))),
4164 x86_64::Piece::Implicit { reg } => match bound(list, reg, desc.role) {
4165 Some(index) => (index, None),
4166 None => return self.spare(inst, desc),
4167 },
4168 // A register the template named, which belongs to one of the statement's operands when
4169 // a constraint letter put that operand there and to nobody otherwise. Asked in that
4170 // order rather than placed straight away, because `"D" (p)` with `%rdi` in the text is
4171 // the program saying one thing twice, and answering it twice would hand the allocator
4172 // one register holding two values.
4173 x86_64::Piece::Reg { reg, .. } => match bound(list, reg, desc.role) {
4174 Some(index) => (index, None),
4175 None => return self.itself(inst, desc, reg),
4176 },
4177 };
4178 let operand = list.get(index).copied().ok_or_else(refused)?;
4179 // The two halves of an operand written `+`, which arrives in one register and leaves in
4180 // another with the allocator told to make them the same one. Everything else has one of
4181 // the two and asking for the other is the refusal below.
4182 let place = places.get(index).copied().ok_or_else(refused)?;
4183 let reg = match desc.role {
4184 Role::Use => place.read,
4185 Role::Def | Role::EarlyDef => place.write,
4186 }
4187 .ok_or_else(refused)?;
4188
4189 // Read where the opcode reads and written where it writes, which is what the first half of
4190 // this asks. An output has a result and an input has a value, an output written `+` has
4191 // both because it is read before it is written, and an output a matching constraint names
4192 // is read as the input that named it. See [`read_as`].
4193 // An output with neither is read as well, and what it holds there is undefined, which
4194 // [`Self::assembly`] says why and puts a zero in a register for.
4195 let placeable = match desc.role {
4196 Role::Use => read_as(list, index).is_some() || operand.result.is_some(),
4197 Role::Def | Role::EarlyDef => operand.result.is_some(),
4198 };
4199 let ty = match (operand.result, operand.value) {
4200 (Some(result), _) => self.source[result].ty,
4201 (None, Some(value)) => self.source[value].ty,
4202 (None, None) => return Err(refused()),
4203 };
4204 let bits = held_bits(ty);
4205 if !placeable || self.class_of(ty) != desc.class {
4206 return Err(refused());
4207 }
4208 if let Some((width, stated)) = spelled {
4209 // An operand the template wrote a width on may be written by an instruction that fills
4210 // more of the register than the object in it does, and the object is then the low part
4211 // of what was written. That is what gmp asks for when it counts the low zero bits of a
4212 // limb into an `unsigned` and spells the count `%q0`: one quadword instruction writes
4213 // the whole register and the `unsigned` is the bottom of it, which is every bit of an
4214 // answer that cannot exceed sixty four anyway.
4215 //
4216 // An operand read at a width the template wrote is the other way round: the object is
4217 // in the register and the instruction looks at the bottom of it. tcc tests the low bits
4218 // of a `size_t` count with `testb $2,%b4`, and every bit that test reads is one the
4219 // object put there.
4220 //
4221 // A write of less of a register than the object fills is right in one case, which is
4222 // an instruction that reads the register it writes and an operand that arrives with
4223 // the object in it. The top of the register is then the top of the object, and the
4224 // instruction leaves it alone. tcc swaps the bytes of an `unsigned` with `xchgb
4225 // %b0,%h0` and a rotate between two of them, and the swap only ever touches the low
4226 // half.
4227 //
4228 // The two that stay refused are a read of more of a register than its type fills,
4229 // which hands an instruction bits nothing ever put there, and a write of less of one
4230 // that nothing carried the object into, which leaves the top of the object holding
4231 // whatever the register held before. An operand the template left plain is refused
4232 // either way, because what gets spelled for that one is the register at the width of
4233 // its type and no other instruction is the one written down.
4234 let carried = matches!(desc.constraint, Constraint::Reuse(_) | Constraint::Fixed(_))
4235 && read_as(list, index).is_some();
4236 // The other case is the one the machine settles by itself: a write of the low four
4237 // bytes of a register clears the four above them, so a sixty four bit object written
4238 // that way holds the thirty two bit answer and nothing else. tcc loads a word through
4239 // `movl 4(%0),%k0` into a `long` and means exactly that.
4240 let cleared = desc.class == self.gpr && width == x86_64::Width::Long && bits == 64;
4241 let widened = stated && desc.role.is_def() && width.bits() > bits;
4242 let narrowed =
4243 stated && width.bits() < bits && (!desc.role.is_def() || carried || cleared);
4244 if bits != width.bits() && !widened && !narrowed {
4245 return Err(refused());
4246 }
4247 }
4248 // An operand the program pinned is in that register and nowhere else, whatever the opcode
4249 // would have allowed it. That is the whole of what a local register variable asks for, and
4250 // it is the same shape a division already has: the allocator is told the register, puts a
4251 // move in front or behind where it has to, and leaves it out where it does not.
4252 let constraint = match pinned(&operand) {
4253 Some(reg) => Constraint::Fixed(reg),
4254 None => desc.constraint,
4255 };
4256 Ok(mir::Operand { reg, class: desc.class, role: desc.role, constraint })
4257 }
4258
4259 /// A register the template named in its own text.
4260 ///
4261 /// Not one of the statement's operands and not something the allocator handed out. The program
4262 /// wrote `%rbx` in the middle of a template and meant that register, which is what code doing
4263 /// something the constraint letters cannot say is made of: micropython saves the callee-saved
4264 /// registers into a buffer by name because the whole point of the buffer is that those exact
4265 /// registers are in it, and there is no constraint letter for `%rsp`.
4266 ///
4267 /// So it is placed as itself, fixed to the register the template named. What that buys is the
4268 /// thing gcc does not do: the register becomes part of the instruction the allocator sees, so a
4269 /// write of one is a definition it knows about and will not leave anything of the program's
4270 /// across, and a read of one is a use it will not have put something else in first. gcc copies
4271 /// the text out and a register two things believe they own is a wrong program nothing reports.
4272 /// Here the allocator is told, and a program that also named the register in its clobber list
4273 /// says the same thing twice rather than something new.
4274 fn itself(
4275 &mut self,
4276 inst: Inst,
4277 desc: OperandDesc,
4278 reg: PhysReg,
4279 ) -> Result<mir::Operand, Unsupported> {
4280 let refused = Unsupported::Assembly { inst, refused: Written::Operand };
4281 if desc.class != self.gpr {
4282 return Err(refused);
4283 }
4284 Ok(mir::Operand {
4285 reg: mir::Reg::physical(reg),
4286 class: self.gpr,
4287 role: desc.role,
4288 constraint: Constraint::Fixed(reg),
4289 })
4290 }
4291
4292 /// A register an instruction of a template uses and the statement put nothing in.
4293 ///
4294 /// A write of one is the register being destroyed, which is what a clobber list is usually
4295 /// written to say and what an instruction with more answers than the program asked for does
4296 /// anyway: `cpuid` writes all four registers whether or not the statement wanted all four. A
4297 /// register of its own is the whole of what that needs, since a value nothing reads is one the
4298 /// allocator may put anywhere and is told about so that nothing else is put there.
4299 ///
4300 /// A read of one is a register the instruction looks at and the program never filled, which
4301 /// gcc leaves as whatever happened to be there. A zero is written instead, for the reason
4302 /// [`Self::undefined`] gives: the allocator has to be given a definition before a use, and a
4303 /// zero is the one answer that reads the same on every run.
4304 fn spare(&mut self, inst: Inst, desc: OperandDesc) -> Result<mir::Operand, Unsupported> {
4305 let refused = Unsupported::Assembly { inst, refused: Written::Operand };
4306 if desc.class != self.gpr {
4307 return Err(refused);
4308 }
4309 let reg = self.out.new_vreg(desc.class);
4310 if !desc.role.is_def() {
4311 let block = self.at.expect("a block is being filled");
4312 let span = self.source.span(inst);
4313 let put = mir::Opcode::new(self.names.intern(&format!("{PREFIX}mov_ri_64")));
4314 self.out.build(block, put).at(span).def(reg, desc.class).imm(0).finish();
4315 }
4316 Ok(mir::Operand { reg, class: desc.class, role: desc.role, constraint: desc.constraint })
4317 }
4318
4319 /// The address one instruction of a template reads or writes.
4320 fn addressed(
4321 &mut self,
4322 inst: Inst,
4323 at: x86_64::At,
4324 places: &[Place],
4325 list: &[AsmOperand<'_>],
4326 ) -> Result<mir::Mem, Unsupported> {
4327 let refused = || Unsupported::Assembly { inst, refused: Written::Operand };
4328 let base = match at.base {
4329 None => None,
4330 Some(x86_64::Piece::Operand { index, .. }) => {
4331 // The register an address is counted from is read and never written, whatever the
4332 // instruction does to what it finds there.
4333 let reg = places.get(index).and_then(|place| place.read).ok_or_else(refused)?;
4334 Some(mir::Operand::read(reg, self.gpr))
4335 }
4336 // A register the template named, counted from as itself. See [`Self::itself`], and note
4337 // that this is the half of it every one of these templates needs: `movq %rax, 16(%rdi)`
4338 // names one register as the thing being stored and another as where to store it. An
4339 // operand a constraint letter put in that register is that operand, for the reason
4340 // [`Self::placed`] gives.
4341 Some(x86_64::Piece::Reg { reg, .. }) => match bound(list, reg, Role::Use) {
4342 Some(index) => {
4343 let reg = places.get(index).and_then(|place| place.read).ok_or_else(refused)?;
4344 Some(mir::Operand::read(reg, self.gpr))
4345 }
4346 None => Some(
4347 mir::Operand::read(mir::Reg::physical(reg), self.gpr)
4348 .with(Constraint::Fixed(reg)),
4349 ),
4350 },
4351 // An address counted from a register the instruction reaches without being told is
4352 // not something this machine has: every addressing mode is written out in the text it
4353 // is part of, so a base that got here another way is a base nothing wrote down.
4354 Some(x86_64::Piece::Implicit { .. }) => return Err(refused()),
4355 };
4356 // A distance the template wrote, or the one in an operand the template pointed at, which is
4357 // the same distance said by something that knows how big a thing is. It has to be a number
4358 // the compiler can read at translation time, since it goes in the instruction rather than
4359 // in a register, and an operand holding anything else is refused rather than put somewhere.
4360 let disp = match at.disp {
4361 x86_64::Disp::Number(disp) => disp,
4362 x86_64::Disp::Operand(index) => {
4363 let value =
4364 list.get(index).and_then(|operand| operand.value).ok_or_else(refused)?;
4365 let number = self.number(value).ok_or_else(refused)?;
4366 i32::try_from(number).map_err(|_| refused())?
4367 }
4368 };
4369 Ok(mir::Mem { base, scale: 1, disp, segment: at.segment, ..mir::Mem::default() })
4370 }
4371
4372 /// The number in that value, for one an `iconst` defined, read at the width of its own type.
4373 ///
4374 /// Signed, because the two things a template asks this for are a distance into an address and
4375 /// the number on an instruction, and both of those are signed wherever they land. A constant
4376 /// whose type is unsigned and whose top bit is set therefore reads as a negative number here,
4377 /// which is the same number and is the reading that fits in the thirty two bits an addressing
4378 /// mode has room for.
4379 fn number(&self, value: Value) -> Option<i128> {
4380 let Def::Result { inst, .. } = self.source[value].def else { return None };
4381 if self.source[inst].opcode != Opcode::IConst {
4382 return None;
4383 }
4384 let Extra::Imm(imm) = self.source[inst].extra else { return None };
4385 let bits = self.source[imm].bits();
4386 let width = self.source[value].ty.bits();
4387 if width == 0 || width > 128 {
4388 return None;
4389 }
4390 let spare = 128 - width;
4391 Some(((bits << spare) as i128) >> spare)
4392 }
4393
4394 /// A register holding a value the program has no claim on, written as a zero.
4395 ///
4396 /// Every other way of saying it costs the same instruction or needs a word the machine IR does
4397 /// not have, and a zero is the one that reads the same on every run.
4398 fn undefined(&mut self, inst: Inst, result: Value) -> Result<(), Unsupported> {
4399 let ty = self.source[result].ty;
4400 let refused = Unsupported::Assembly { inst, refused: Written::Operand };
4401 let bits = held_bits(ty);
4402 if self.class_of(ty) != self.gpr || !matches!(bits, 8 | 16 | 32 | 64) {
4403 return Err(refused);
4404 }
4405 let block = self.at.expect("a block is being filled");
4406 let span = self.source.span(inst);
4407 let reg = self.new_reg(result);
4408 let put = mir::Opcode::new(self.names.intern(&format!("{PREFIX}mov_ri_{bits}")));
4409 self.out.build(block, put).at(span).def(reg, self.gpr).imm(0).finish();
4410 Ok(())
4411 }
4412
4413 /// Whether a type is the width an address is, which is what makes a cast to or from one free.
4414 fn is_address_width(&self, ty: Type) -> bool {
4415 ty.is_ptr() || (ty.is_int() && ty.bits() == ADDRESS_BITS)
4416 }
4417
4418 /// Where a block goes, which in machine IR is on the block rather than on its terminator.
4419 ///
4420 /// That is why no rule ever names a block: a branch is selected for what it reads and the
4421 /// edges are copied across here, arguments and all. The arguments are read last, after every
4422 /// instruction of the block is written, because an argument that is a constant is
4423 /// materialized where it is first wanted and the end of the block is where an edge wants it.
4424 ///
4425 /// Which is not quite the end. A block that leaves two ways has the branch as its last
4426 /// instruction, and a block that leaves through a register has the indirect jump as its last,
4427 /// and anything appended after either is something it has already jumped past, so a constant
4428 /// materialized here would be a register the block below reads and nothing ever writes. The
4429 /// one that was there is put back on the end when that happened, which is the only reordering
4430 /// anything in this crate does and is why it is remembered before a single argument is read.
4431 fn edges(&mut self, block: Block, out: mir::Block) -> Result<(), Unsupported> {
4432 let Some(term) = self.source.terminator(block) else { return Ok(()) };
4433 let leaves =
4434 matches!(self.source[term].opcode, Opcode::BrIf | Opcode::IndirectBr | Opcode::Switch);
4435 let branch = if leaves { self.out.terminator(out) } else { None };
4436
4437 let calls: Vec<rucc_ir::BlockCall> = self.source.successors(term).collect();
4438 let mut succs = Vec::with_capacity(calls.len());
4439 for call in calls {
4440 let args: Vec<Value> = self.source[call.args].to_vec();
4441 let mut regs = Vec::with_capacity(args.len());
4442 for value in args {
4443 // The address of where the value is rather than the value, for the one type a
4444 // register holds none of. The block on the other side copies the bytes out of it
4445 // into a slot of its own, which is what makes a second edge into the same block
4446 // safe.
4447 let reg = if on_x87(self.source[value].ty) {
4448 self.x87_slot(value)
4449 } else {
4450 self.reg_of(value)?
4451 };
4452 regs.push(reg);
4453 }
4454 succs.push(mir::BlockCall::with(self.out_block(call.block), regs));
4455 }
4456 if let Some(branch) = branch {
4457 if self.out.terminator(out) != Some(branch) {
4458 self.out.remove_inst(branch);
4459 self.out.append_inst(out, branch);
4460 }
4461 }
4462 *self.out.succs_mut(out) = succs;
4463 Ok(())
4464 }
4465
4466 /// The machine IR block an IR block became.
4467 fn out_block(&self, block: Block) -> mir::Block {
4468 self.blocks[block.index()].expect("every block was created before any was filled")
4469 }
4470
4471 /// The parameters of the entry block, which are the function's arguments.
4472 ///
4473 /// They are not block parameters in the machine IR and they cannot be. A block parameter is
4474 /// given its value by a move on the edge into the block, and there is no edge into an entry
4475 /// block, so what arrives in a function is the convention's to say. [`crate::abi`] is what
4476 /// says it.
4477 ///
4478 /// The ones past the last register arrived in the caller's memory and are read out of it, and
4479 /// the loads that read them come back here so that the frame can finish them the way it
4480 /// finishes an `alloca`.
4481 fn arrive(&mut self, block: Block, out: mir::Block) -> Result<(), Unsupported> {
4482 let params = self.source[block].params.clone();
4483 // The type of each is the block's answer and what the ABI asks of it is the signature's,
4484 // and the two lists are the same list: a parameter the classification turned into a
4485 // pointer is a pointer in the block too. A block with more parameters than the signature
4486 // names is not one the front end writes, and each of those is taken as a plain value.
4487 let asked: Vec<Abi> = self.source.signature().params.iter().map(|it| it.abi).collect();
4488 let types: Vec<Param> = params
4489 .iter()
4490 .enumerate()
4491 .map(|(index, &value)| {
4492 let abi = asked.get(index).copied().unwrap_or_default();
4493 Param { ty: self.source[value].ty, abi }
4494 })
4495 .collect();
4496 // A save area for a function that takes arguments its signature does not name, which is a
4497 // block of this function's frame on one convention and the shadow space the caller already
4498 // reserved on the other. Which of the two it is is [`varargs::Area::of`]'s answer and
4499 // [`Self::save_area`] is where the difference is spent.
4500 let variadic = self.source.signature().variadic;
4501 let area = variadic.then(|| varargs::Area::of(self.conv));
4502 let arrived = abi::entry(&mut self.out, out, &types, self.conv, self.names, area)
4503 .map_err(|(index, missing)| Unsupported::Argument { index, missing })?;
4504 for (¶m, reg) in params.iter().zip(&arrived.regs) {
4505 self.regs[param.index()] = Some(*reg);
4506 }
4507 if let Some(area) = area {
4508 self.save_area(out, &arrived, area);
4509 }
4510 self.stack.arguments.extend(arrived.stack);
4511 Ok(())
4512 }
4513
4514 /// The prologue of a variadic function, which is every argument register it was handed written
4515 /// into the frame.
4516 ///
4517 /// Every one the signature did not name, that is. Which of those hold anything is a thing only
4518 /// the caller knew and there is nothing here to ask, so all of them are written, and the ones a
4519 /// named parameter took are not, because `va_start` sets the two offsets past them and nothing
4520 /// ever reads their slots.
4521 ///
4522 /// What that costs is up to fourteen stores in the prologue of a function that may read none of
4523 /// them, and the convention's answer to that is the count of vector registers in `%al`, which
4524 /// lets a callee skip the eight vector stores when the call passed no floats. Skipping them is a
4525 /// branch in a prologue, and a prologue is written long after this by [`crate::finish`], which
4526 /// has no blocks to branch between. So they are all written every time, which is correct and is
4527 /// what `-O0` costs. Issue #323 is the branch.
4528 ///
4529 /// A vector register is written all sixteen bytes at a time, because a `_Float128` fills one and
4530 /// a `va_arg` of a quad reads the slot back whole. gcc writes the same sixteen with the same
4531 /// instruction, which is what [`crate::varargs`] says a list has to be built out of.
4532 ///
4533 /// The address is computed once into a register rather than written as a displacement off the
4534 /// stack pointer, because a displacement into a frame is not known until after allocation and
4535 /// one `lea` costs less than a fixup list for a dozen stores. It is the same `lea` an `alloca`
4536 /// gets and [`crate::finish`] fills it in the same way.
4537 ///
4538 /// A convention that homes its register arguments has none of that. Its area is the shadow
4539 /// space the caller reserved above the return address, so there is no object to make and no
4540 /// address to work out: each store reaches into the caller's argument area the way the load of
4541 /// a parameter the registers ran out before does, which is the same waiting list and the same
4542 /// fixup. There are at most four of them and none is a vector register, since a float the
4543 /// signature does not name arrived in a general purpose register too and that is the copy the
4544 /// walk reads.
4545 fn save_area(&mut self, out: mir::Block, arrived: &abi::Arrived, area: varargs::Area) {
4546 if self.conv.shared_positions {
4547 self.varargs = Some(Varargs::Pointer { incoming: arrived.beyond });
4548 let store = mir::Opcode::new(self.names.intern("x64.mov_mr_64"));
4549 for &(reg, class, at) in &arrived.spare {
4550 let sp = mir::Operand::read(mir::Reg::physical(self.conv.stack_pointer), self.gpr);
4551 let made =
4552 self.out.build(out, store).uses(reg, class).mem(mir::Mem::at(sp)).finish();
4553 self.stack.arguments.push((made, at));
4554 }
4555 return;
4556 }
4557
4558 let save = self.stack.locals.len();
4559 self.stack.locals.push(Local { size: area.size, align: varargs::VECTOR_SLOT });
4560 self.varargs = Some(Varargs::Fields {
4561 save,
4562 incoming: arrived.beyond,
4563 integers: u32::try_from(arrived.took.0).unwrap_or(0) * area.stride(false),
4564 floats: area.starts_at(true)
4565 + u32::try_from(arrived.took.1).unwrap_or(0) * area.stride(true),
4566 });
4567
4568 let base = self.frame_address(out, save);
4569 for &(reg, class, at) in &arrived.spare {
4570 let name = if class == self.gpr { "x64.mov_mr_64" } else { "x64.movaps_mr" };
4571 let store = mir::Opcode::new(self.names.intern(name));
4572 let up = i32::try_from(at).expect("a register save area under two gigabytes");
4573 let mem = mir::Mem::at(mir::Operand::read(base, self.gpr)).plus(up);
4574 self.out.build(out, store).uses(reg, class).mem(mem).finish();
4575 }
4576 }
4577
4578 /// The address of one of the function's stack objects, in a fresh register.
4579 ///
4580 /// Written with nothing in its displacement, because where an object is in a frame is not known
4581 /// until after allocation, and given to [`crate::finish`] to fill in the way an `alloca` is.
4582 fn frame_address(&mut self, out: mir::Block, local: usize) -> mir::Reg {
4583 let reg = self.out.new_vreg(self.gpr);
4584 let lea = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.lea)));
4585 let sp = mir::Operand::read(mir::Reg::physical(self.conv.stack_pointer), self.gpr);
4586 let made = self.out.build(out, lea).def(reg, self.gpr).mem(mir::Mem::at(sp)).finish();
4587 self.stack.addresses.push((made, local));
4588 reg
4589 }
4590
4591 /// Whether an instruction is one no machine instruction is written for where it stands.
4592 ///
4593 /// Four of them, and none is a lowering decision, which is why none is a rule. A constant is
4594 /// written where a register for it is first wanted rather than where the IR put it, and every
4595 /// reader of one may have folded it into an immediate, in which case nowhere is the right
4596 /// place. A return of nothing has nothing to put anywhere: the epilogue gives the frame back
4597 /// and leaves, and it is appended to every block with no successors long after this has
4598 /// finished, so a return with a value is one instruction here and a return without one is
4599 /// none. Unless the value went back through memory, in which case there is something to put
4600 /// somewhere after all and the IR does not carry it: the address the caller handed over has
4601 /// to be in `rax` on the way out, and [`Lowering::returned`] is what writes that.
4602 ///
4603 /// An unconditional jump is the third, and there is even less of it: the edge is on the
4604 /// block, and whether the block it goes to is the next one and needs no jump at all is the
4605 /// block layout's answer rather than this one's.
4606 ///
4607 /// The fourth is a point control does not arrive at, in both of the forms the IR has for it:
4608 /// the `unreachable` terminator the front end puts at the end of a function whose body can run
4609 /// off the bottom, and the `unreachable_hint` a call to `__builtin_unreachable` becomes. What
4610 /// to write for a place nothing reaches is a question with no wrong answer, and nothing is the
4611 /// smallest one and the one gcc 16.2.0 gives at `-O0`. The terminator leaves the block with no
4612 /// successors, so the epilogue lands at the end of it the way it does on any other block that
4613 /// goes nowhere, and the function cannot fall out of its own last instruction into whatever
4614 /// the assembler puts next.
4615 fn writes_nothing(&self, inst: Inst) -> bool {
4616 let data = &self.source[inst];
4617 match data.opcode {
4618 Opcode::IConst | Opcode::Jump | Opcode::Unreachable | Opcode::UnreachableHint => true,
4619 Opcode::Return => self.source[data.args].is_empty() && self.sret().is_none(),
4620 _ => false,
4621 }
4622 }
4623
4624 /// What every instruction in one block matched, with a set of values nobody may take.
4625 ///
4626 /// Backwards, because an instruction that has been folded into a later one does not get to
4627 /// fold anything into itself: the rule that took it only reached one level down, so what is
4628 /// under it is not in the term the matcher saw and cannot be replaced.
4629 fn decide(&self, insts: &[Inst], refused: &HashSet<Value>) -> Decided {
4630 let mut found: Vec<Option<Match<Term>>> = (0..insts.len()).map(|_| None).collect();
4631 let mut plans: Vec<Option<Plan>> = vec![None; insts.len()];
4632 let mut folded: Vec<Inst> = Vec::new();
4633 for (index, &inst) in insts.iter().enumerate().rev() {
4634 if folded.contains(&inst) {
4635 continue;
4636 }
4637 if let Some((plan, matched)) = self.select(inst, refused) {
4638 folded.extend(self.folds(inst, plan));
4639 found[index] = Some(matched);
4640 plans[index] = Some(plan);
4641 }
4642 }
4643 Decided { found, plans, folded }
4644 }
4645
4646 /// A value some of its readers took and some of them did not, which is the one case folding
4647 /// buys nothing.
4648 ///
4649 /// Folding does not delete the instruction that computed a value for anybody else, so a
4650 /// reader that did not take it still needs it in a register and the instruction stays. The
4651 /// reader that did take it now does that work again. Either all of them take it, in which
4652 /// case nothing is left to read it and the instruction goes, or none of them do.
4653 ///
4654 /// The count is over the whole function rather than over the block, since a value read from
4655 /// another block is read from a register there whatever this block decides. An instruction
4656 /// built by name rather than matched, a call being the one that matters, has no plan and so
4657 /// takes nothing, which is the right answer for it as well.
4658 fn left_alive(&self, insts: &[Inst], plans: &[Option<Plan>]) -> Option<Value> {
4659 let mut taken = vec![0u32; self.uses.len()];
4660 for (&inst, plan) in insts.iter().zip(plans) {
4661 let Some(plan) = plan else { continue };
4662 let args = &self.source[self.source[inst].args];
4663 for (index, &arg) in args.iter().take(MAX_ARGS).enumerate() {
4664 if plan[index] == Shown::Expand {
4665 taken[arg.index()] += 1;
4666 }
4667 }
4668 }
4669 for (&inst, plan) in insts.iter().zip(plans) {
4670 let Some(plan) = plan else { continue };
4671 let args = &self.source[self.source[inst].args];
4672 for (index, &arg) in args.iter().take(MAX_ARGS).enumerate() {
4673 if plan[index] == Shown::Expand && taken[arg.index()] < self.uses[arg.index()] {
4674 return Some(arg);
4675 }
4676 }
4677 }
4678 None
4679 }
4680
4681 /// The rule that fires on an instruction, and what it bound.
4682 ///
4683 /// The plans are tried in order and the first that matches wins, which is the maximal munch
4684 /// `spec/10-backend.md` asks for: a plan that offers more to the matcher is tried before one
4685 /// that offers less.
4686 fn select(&self, inst: Inst, refused: &HashSet<Value>) -> Option<(Plan, Match<Term>)> {
4687 for plan in self.plans(inst, refused) {
4688 let terms = Terms::new(self.source, inst, plan);
4689 if let Some(matched) = TABLE.find(&terms, Term::Root) {
4690 return Some((plan, matched));
4691 }
4692 }
4693 None
4694 }
4695
4696 /// Every way this instruction can be shown to the matcher, most offered first.
4697 fn plans(&self, inst: Inst, refused: &HashSet<Value>) -> Vec<Plan> {
4698 let args = &self.source[self.source[inst].args];
4699 let mut plans = vec![PLAIN];
4700 for (index, &arg) in args.iter().enumerate().take(MAX_ARGS) {
4701 let mut ways = Vec::new();
4702 if self.foldable(inst, arg, refused) {
4703 ways.push(Shown::Expand);
4704 }
4705 if Terms::new(self.source, inst, PLAIN).constant(arg).is_some() {
4706 ways.push(Shown::Const);
4707 }
4708 ways.push(Shown::Reg);
4709 plans = plans
4710 .into_iter()
4711 .flat_map(|plan| {
4712 ways.iter().map(move |&way| {
4713 let mut next = plan;
4714 next[index] = way;
4715 next
4716 })
4717 })
4718 .collect();
4719 }
4720 plans
4721 }
4722
4723 /// Whether an operand may be shown as the instruction that computed it.
4724 ///
4725 /// It has to be in the same block, because a rule that folds one instruction into another
4726 /// moves the work to where the second one is. It has to be something rather than a block
4727 /// parameter, and not a constant, which is shown as a constant instead. And it has to be a
4728 /// value [`Lowering::left_alive`] has not put back, which is how the one reader at a time
4729 /// question is asked here: this says yes to a value with any number of readers, and a value
4730 /// only some of them could take is refused after the fact and asked again.
4731 ///
4732 /// A value with several readers used to be refused outright, on the reasoning that folding
4733 /// does not delete the instruction for anybody else. That reasoning is about the set of
4734 /// readers and was being applied to one reader at a time, which is stricter than it needs to
4735 /// be: when every reader takes it there is nobody left to read it and the instruction goes.
4736 /// An address a store and a load share is the shape that matters, since a memory operand has
4737 /// room for the whole of it and both readers have a memory operand.
4738 fn foldable(&self, into: Inst, value: Value, refused: &HashSet<Value>) -> bool {
4739 let Def::Result { inst, .. } = self.source[value].def else { return false };
4740 if self.source[inst].opcode == Opcode::IConst || refused.contains(&value) {
4741 return false;
4742 }
4743 self.source.block_of(inst).is_some()
4744 && self.source.block_of(inst) == self.source.block_of(into)
4745 }
4746
4747 /// The instructions a match folded into the one it matched.
4748 ///
4749 /// The plan is what says this, not the bindings: a binding is a register or a number either
4750 /// way, and an operand shown as the instruction that computed it is one no rule could have
4751 /// matched without taking that instruction, because the plan offered the matcher nothing
4752 /// else to call it.
4753 fn folds(&self, inst: Inst, plan: Plan) -> Vec<Inst> {
4754 let args = &self.source[self.source[inst].args];
4755 args.iter()
4756 .take(MAX_ARGS)
4757 .enumerate()
4758 .filter(|&(index, _)| plan[index] == Shown::Expand)
4759 .filter_map(|(_, &arg)| match self.source[arg].def {
4760 Def::Result { inst, .. } => Some(inst),
4761 Def::Param { .. } => None,
4762 })
4763 .collect()
4764 }
4765
4766 /// What the IR instruction said about itself that the machine instruction has to keep saying.
4767 ///
4768 /// One flag today. `volatile` says the access happens exactly once and is never moved or
4769 /// merged, and nothing below here can work that out again: a `volatile` load and an ordinary
4770 /// one are the same instruction over the same address, so a pass that puts two accesses
4771 /// together would put these together too. Carried rather than checked here, because the pass
4772 /// that has to refuse is a long way down and this is the last place the answer is known.
4773 ///
4774 /// The instructions this compiler writes for itself get nothing, which is the right answer
4775 /// for all of them: a prologue, a spill and the moves around a call were asked for by the
4776 /// machine rather than by the program.
4777 ///
4778 /// Every access the flag is legal on carries it: the loads and the stores a rule matched,
4779 /// the two ends of a `long double` copy that are the program's own memory, and the compare
4780 /// and exchange and the read modify write. An `asm` statement does not, and it is the one
4781 /// exception on purpose. What the flag says there is that the statement stays even when
4782 /// nothing reads what it wrote, which is a different sentence about a different thing, and
4783 /// every `asm` is already fixed where it stands whether the word was written or not.
4784 fn carried(&self, inst: Inst) -> mir::Flags {
4785 if self.source[inst].flags.contains(Flags::VOLATILE) {
4786 mir::Flags::VOLATILE
4787 } else {
4788 mir::Flags::NONE
4789 }
4790 }
4791
4792 /// Build the machine instruction a match calls for.
4793 fn emit(&mut self, inst: Inst, matched: &Match<Term>) -> Result<(), Unsupported> {
4794 let rule: &Rule = TABLE.rule(matched);
4795 let pieces = rule.replacement;
4796 let Some(Piece::App { head, arity }) = pieces.first() else {
4797 return Err(self.unsupported(inst));
4798 };
4799 let opcode = head.strip_prefix(PREFIX).ok_or_else(|| self.unsupported(inst))?;
4800 let form = x86_64::form(opcode).ok_or_else(|| self.unsupported(inst))?;
4801
4802 let mut read = Read::default();
4803 let mut at = 1;
4804 for _ in 0..*arity {
4805 at = self.read(inst, pieces, at, &matched.bindings, &mut read)?;
4806 }
4807
4808 let descs = form.operands();
4809 let writes = descs.iter().take_while(|desc| desc.role.is_def()).count();
4810 if descs.len() - writes != read.regs.len() {
4811 return Err(self.unsupported(inst));
4812 }
4813
4814 // The first thing the instruction writes is what it computes, and any others are
4815 // registers the machine destroys on the way, which are fresh because nothing else is in
4816 // them and nothing reads them. An instruction that writes nothing at all is one whose
4817 // whole purpose is its effect, which is what a store is, and there is no result to put
4818 // anywhere.
4819 let mut regs = Vec::new();
4820 if writes > 0 {
4821 let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
4822 regs.push(self.new_reg(result));
4823 // The rest are the registers the machine destroys on the way, and the class each is in
4824 // is the one the instruction's description gives it rather than a guess, so that an
4825 // instruction that wrecks a register in the other file says so.
4826 regs.extend(descs[1..writes].iter().map(|desc| self.out.new_vreg(desc.class)));
4827 } else if self.source[inst].first_result.is_some() {
4828 // A rule that throws away a value the IR gave a name to would leave every reader of
4829 // that name with nothing to read, so it is a rule this and the target disagree about.
4830 return Err(self.unsupported(inst));
4831 }
4832 regs.extend(read.regs.iter().copied());
4833
4834 let block = self.at.expect("a block is being filled");
4835 let opcode = mir::Opcode::new(self.names.intern(head));
4836 let (span, flags) = (self.source.span(inst), self.carried(inst));
4837 let mut build = self.out.build(block, opcode).at(span).flags(flags);
4838 for (desc, reg) in descs.iter().zip(regs) {
4839 let operand = mir::Operand {
4840 reg,
4841 class: desc.class,
4842 role: desc.role,
4843 constraint: desc.constraint,
4844 };
4845 build = build.operand(operand);
4846 }
4847 if let Some(mem) = read.mem {
4848 build = build.mem(mem);
4849 }
4850 if let Some(imm) = read.imm {
4851 build = build.imm(imm);
4852 }
4853 build.finish();
4854 Ok(())
4855 }
4856
4857 /// Read one argument of a replacement, which is a register, a number or an address.
4858 ///
4859 /// Gives back the position after it, because a replacement is flat and an address takes
4860 /// arguments of its own.
4861 fn read(
4862 &mut self,
4863 inst: Inst,
4864 pieces: &'static [Piece],
4865 at: usize,
4866 bindings: &[Term],
4867 out: &mut Read,
4868 ) -> Result<usize, Unsupported> {
4869 match pieces.get(at) {
4870 Some(Piece::Int(value)) => {
4871 out.imm = i64::try_from(*value).ok();
4872 Ok(at + 1)
4873 }
4874 // A number the rule worked out of the ones it matched rather than one it wrote down,
4875 // which is an immediate once it has been worked out and is read here as one. It gives
4876 // nothing back when a binding it reads is a register, and a replacement that cannot be
4877 // built is a rule this file and the matcher disagree about, which is what `unsupported`
4878 // is for.
4879 Some(Piece::Computed { work, .. }) => {
4880 let matched: Vec<Option<i128>> = bindings
4881 .iter()
4882 .map(|term| match *term {
4883 Term::Num(value) => Some(value),
4884 _ => None,
4885 })
4886 .collect();
4887 let number = work(&matched).ok_or_else(|| self.unsupported(inst))?;
4888 out.imm = i64::try_from(number).ok();
4889 Ok(at + 1)
4890 }
4891 Some(Piece::Var { index, .. }) => {
4892 match bindings.get(*index) {
4893 Some(&Term::Reg(value)) => {
4894 let reg = self.reg_of(value)?;
4895 out.regs.push(reg);
4896 }
4897 Some(&Term::Num(value)) => out.imm = i64::try_from(value).ok(),
4898 // A pattern binds a register or a number and nothing else, so this is a
4899 // rule the matcher and this file disagree about.
4900 _ => return Err(self.unsupported(inst)),
4901 }
4902 Ok(at + 1)
4903 }
4904 Some(Piece::App { head, arity }) => {
4905 let kind = x86_64::address(head).ok_or_else(|| self.unsupported(inst))?;
4906 let mut inner = Read::default();
4907 let mut next = at + 1;
4908 for _ in 0..*arity {
4909 next = self.read(inst, pieces, next, bindings, &mut inner)?;
4910 }
4911 let mem = address(kind, &inner, self.gpr).ok_or_else(|| self.unsupported(inst))?;
4912 out.mem = Some(mem);
4913 Ok(next)
4914 }
4915 None => Err(self.unsupported(inst)),
4916 }
4917 }
4918
4919 /// The register a value is in, materializing it if it is a constant that has not been put in
4920 /// one yet.
4921 ///
4922 /// A constant is written where it is wanted rather than where the IR defined it, and where it
4923 /// is wanted is a block that need not be the one the IR defined it in. So the register holding
4924 /// one is only good inside the block it was written into, and a second block that wants the
4925 /// same constant gets its own. Anything else is a register read where nothing wrote it: the
4926 /// IR guarantees a definition dominates its uses, and this moved the definition.
4927 ///
4928 /// Writing the number again is also the right answer and not merely the safe one. It is one
4929 /// instruction that reads nothing, which is cheaper than holding a register live across a
4930 /// branch for it, and it is what a rematerializing allocator would do with the value anyway.
4931 fn reg_of(&mut self, value: Value) -> Result<mir::Reg, Unsupported> {
4932 let constant = match self.source[value].def {
4933 Def::Result { inst, .. } => {
4934 (self.source[inst].opcode == Opcode::IConst).then_some(inst)
4935 }
4936 Def::Param { .. } => None,
4937 };
4938 let here = self.at.expect("a block is being filled");
4939 if let Some(reg) = self.regs[value.index()] {
4940 if constant.is_none() || self.written[value.index()] == Some(here) {
4941 return Ok(reg);
4942 }
4943 }
4944 if let Some(inst) = constant {
4945 // Cleared so that the register the constant is written into is a new one rather than
4946 // the one the block above wrote, which is still being read up there.
4947 self.regs[value.index()] = None;
4948 // Nothing is refused here. A constant is written on its own, out of the loop over the
4949 // block, and the operands of the rule that writes one are the number and nothing else.
4950 let matched = self
4951 .select(inst, &HashSet::new())
4952 .map(|(_, matched)| matched)
4953 .ok_or_else(|| self.unsupported(inst))?;
4954 self.emit(inst, &matched)?;
4955 // The same mark the loop over the instructions makes, and it has to be made here as
4956 // well because this is the only place a constant is ever selected: the loop skips one
4957 // where the IR wrote it, so a rule that lowers a constant fires from nowhere else and
4958 // would be reported as a rule nothing reaches.
4959 self.fired.mark(matched.rule);
4960 self.written[value.index()] = Some(here);
4961 return Ok(self.regs[value.index()].expect("a constant is written into a register"));
4962 }
4963 Ok(self.new_reg(value))
4964 }
4965
4966 /// Which register file a value of that type lives in.
4967 ///
4968 /// The vector one for the two float widths the machine has scalar instructions for and for the
4969 /// one it only moves, and the general purpose one for everything else. An eighty bit `long
4970 /// double` is in neither, and it is here rather than in the vector class on purpose: it would
4971 /// be put in a register that cannot hold it, and there is no rule that names one, so the
4972 /// instruction computing it is reported. The wrong class would make that a wrong program
4973 /// instead of a refused one.
4974 ///
4975 /// A hundred and twenty eight bit float is in the vector class and fits it exactly, which is
4976 /// the difference. Nothing computes in it, so every arithmetic on one is still reported, and
4977 /// what the class buys is the moves: a register that holds the whole value is a register a
4978 /// spill, a reload and a copy are each one instruction for.
4979 fn class_of(&self, ty: Type) -> RegClass {
4980 if crate::term::in_vector_file(ty) { self.conv.sse_class } else { self.gpr }
4981 }
4982
4983 /// A fresh register for a value, which is what the instruction computing it writes.
4984 ///
4985 /// Any declaration the value is a value of comes with it. Here rather than once at the end over
4986 /// the whole map, because a constant is written again in every block that wants one and the map
4987 /// only remembers the last of those registers, and a local held in a constant is a local that
4988 /// would otherwise be findable in one block of the function and nowhere else.
4989 fn new_reg(&mut self, value: Value) -> mir::Reg {
4990 if let Some(reg) = self.regs[value.index()] {
4991 return reg;
4992 }
4993 let reg = self.out.new_vreg(self.class_of(self.source[value].ty));
4994 self.regs[value.index()] = Some(reg);
4995 let source = self.source;
4996 for decl in source.value_decls(value) {
4997 self.out.named.push((decl, reg));
4998 }
4999 reg
5000 }
5001
5002 fn unsupported(&self, inst: Inst) -> Unsupported {
5003 let data = &self.source[inst];
5004 Unsupported::Inst {
5005 inst,
5006 term: Terms::new(self.source, inst, PLAIN).name(inst),
5007 opcode: data.opcode,
5008 ty: data.first_result.map(|result| self.source[result].ty),
5009 }
5010 }
5011}
5012
5013/// What the arguments of one replacement came to.
5014#[derive(Debug, Default)]
5015struct Read {
5016 regs: Vec<mir::Reg>,
5017 imm: Option<i64>,
5018 mem: Option<mir::Mem>,
5019}
5020
5021/// The addressing mode an address constructor's arguments make.
5022///
5023/// One arm per constructor rather than a question asked of the kind, because what the arguments
5024/// mean is the whole of what tells the four apart: the same register is a base in one and an
5025/// index in another, and the same constant is a scale in one and a displacement in another.
5026fn address(kind: x86_64::Address, read: &Read, gpr: RegClass) -> Option<mir::Mem> {
5027 let mut regs = read.regs.iter().copied().map(|reg| mir::Operand::read(reg, gpr));
5028 match kind {
5029 x86_64::Address::BaseIndexScale => {
5030 let base = regs.next()?;
5031 let index = regs.next()?;
5032 Some(mir::Mem::at(base).indexed(index, u8::try_from(read.imm?).ok()?))
5033 }
5034 x86_64::Address::IndexScale => Some(mir::Mem {
5035 base: None,
5036 index: Some(regs.next()?),
5037 scale: u8::try_from(read.imm?).ok()?,
5038 disp: 0,
5039 symbol: None,
5040 block: None,
5041 table: None,
5042 reach: mir::Reach::Itself,
5043 segment: None,
5044 }),
5045 x86_64::Address::Base => Some(mir::Mem::at(regs.next()?)),
5046 // The rule that writes this has a guard saying the constant fits, so a displacement that
5047 // does not is a rule and a target that disagree rather than a program this cannot compile.
5048 x86_64::Address::BaseOffset => {
5049 Some(mir::Mem { disp: i32::try_from(read.imm?).ok()?, ..mir::Mem::at(regs.next()?) })
5050 }
5051 }
5052}
5053
5054/// The table this selector matches with.
5055///
5056/// One target for now, because one target has a rule file. Which table to use becomes a question
5057/// the moment a second one does, and the answer will be the target the session was given rather
5058/// than a constant here.
5059static TABLE: &Table = &crate::select::x86_64::TABLE;
5060
5061#[cfg(test)]
5062mod tests {
5063 use rucc_ir::{
5064 AsmInfo, Builder, CallInfo, Flags, InstData, MemInfo, MemOrder, Restrict, Signature, Type,
5065 };
5066 use rucc_regalloc::assign::Env;
5067 use rucc_target::x86_64::{FRAME, REGS, SYSV};
5068
5069 use super::*;
5070 use crate::finish::{Convention, finish};
5071 use crate::frame::{Frame, Incoming, Layout};
5072
5073 /// A function of as many 64 bit parameters as the test wants, and the block they are in.
5074 fn blank(params: &[Type]) -> (Interner, Func, Block, Vec<Value>) {
5075 let mut names = Interner::new();
5076 let mut func = Func::new(names.intern("f"), Signature::new());
5077 let block = func.create_block();
5078 let values = params.iter().map(|&ty| func.append_param(block, ty)).collect();
5079 (names, func, block, values)
5080 }
5081
5082 /// An ordinary access: not atomic, and aligned enough that nothing here has an opinion.
5083 /// Neither field reaches selection, which is the point of saying it once here.
5084 fn plain() -> MemInfo {
5085 MemInfo {
5086 size: 0,
5087 align: 1,
5088 order: MemOrder::NotAtomic,
5089 tbaa: None,
5090 owns: 0,
5091 restrict: Restrict::NONE,
5092 }
5093 }
5094
5095 /// What the allocator is given: every integer register the convention offers except two, held
5096 /// back so that a move on an edge has somewhere to break a cycle and a spilled value has
5097 /// somewhere to be read into. Which two does not matter, and holding back the last two the
5098 /// convention would reach for leaves every expectation below unchanged.
5099 fn env() -> Env {
5100 const SCRATCH: [PhysReg; 2] = [x86_64::R10, x86_64::R11];
5101 let order: Vec<PhysReg> =
5102 SYSV.int_order.iter().copied().filter(|reg| !SCRATCH.contains(reg)).collect();
5103 Env::new().with(x86_64::GPR, &order, &SCRATCH)
5104 }
5105
5106 /// The machine IR text a function lowers to.
5107 fn lower(names: &mut Interner, source: &Func) -> String {
5108 let out = func(source, names, &SYSV, &Elsewhere::default())
5109 .expect("every instruction has a rule");
5110 mir::print_func(&out.func, names, ®S)
5111 }
5112
5113 #[test]
5114 fn an_addition_of_two_registers_is_one_instruction() {
5115 let i32 = Type::int(32);
5116 let (mut names, mut func, block, args) = blank(&[i32, i32]);
5117 let mut build = Builder::new(&mut func, block);
5118 build.binary(Opcode::Add, args[0], args[1], Flags::default());
5119
5120 assert_eq!(
5121 lower(&mut names, &func),
5122 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
5123 %1:gpr($rsi) = x64.arg_val_32\n %2:gpr(reuse 1) = x64.add_rr_32 %0, %1\n}\n"
5124 );
5125 }
5126
5127 #[test]
5128 fn a_constant_operand_becomes_an_immediate() {
5129 let i32 = Type::int(32);
5130 let (mut names, mut func, block, args) = blank(&[i32]);
5131 let mut build = Builder::new(&mut func, block);
5132 let seven = build.iconst(i32, 7);
5133 build.binary(Opcode::Add, args[0], seven, Flags::default());
5134
5135 // The constant is in the instruction and nothing was written to hold it, which is what
5136 // materializing one where a register for it is wanted buys.
5137 assert_eq!(
5138 lower(&mut names, &func),
5139 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
5140 %1:gpr(reuse 1) = x64.add_ri_32 %0, 7\n}\n"
5141 );
5142 }
5143
5144 #[test]
5145 fn a_constant_too_wide_for_an_immediate_goes_into_a_register() {
5146 let i64 = Type::int(64);
5147 let (mut names, mut func, block, args) = blank(&[i64]);
5148 let mut build = Builder::new(&mut func, block);
5149 let big = build.iconst(i64, i128::from(i32::MAX) + 1);
5150 build.binary(Opcode::Add, args[0], big, Flags::default());
5151
5152 // Nobody wrote this fallback down. The rule that takes an immediate has a guard that
5153 // turns a number this wide down, so it does not fire, and the next way of showing the
5154 // operand puts it in a register.
5155 assert_eq!(
5156 lower(&mut names, &func),
5157 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
5158 %1:gpr = x64.mov_ri_64 2147483648\n %2:gpr(reuse 1) = x64.add_rr_64 %0, %1\n}\n"
5159 );
5160 }
5161
5162 #[test]
5163 fn an_index_calculation_folds_into_an_address() {
5164 let i64 = Type::int(64);
5165 let (mut names, mut func, block, args) = blank(&[i64, i64]);
5166 let mut build = Builder::new(&mut func, block);
5167 let four = build.iconst(i64, 4);
5168 let scaled = build.binary(Opcode::Mul, args[1], four, Flags::default());
5169 build.binary(Opcode::Add, args[0], scaled, Flags::default());
5170
5171 // Three IR instructions and one machine instruction. The multiply is gone because the
5172 // rule that matched reached down and took it.
5173 assert_eq!(
5174 lower(&mut names, &func),
5175 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
5176 %1:gpr($rsi) = x64.arg_val_64\n %2:gpr = x64.lea_64 [%0 + %1*4]\n}\n"
5177 );
5178 }
5179
5180 #[test]
5181 fn an_instruction_every_reader_can_take_is_folded_into_all_of_them() {
5182 let i64 = Type::int(64);
5183 let (mut names, mut func, block, args) = blank(&[i64, i64]);
5184 let mut build = Builder::new(&mut func, block);
5185 let four = build.iconst(i64, 4);
5186 let scaled = build.binary(Opcode::Mul, args[1], four, Flags::default());
5187 let first = build.binary(Opcode::Add, args[0], scaled, Flags::default());
5188 build.binary(Opcode::Add, first, scaled, Flags::default());
5189
5190 // Both readers have room for a scaled index, so both of them take it and nothing is left
5191 // to read the multiply. Three IR instructions become two machine ones, where refusing to
5192 // fold into either reader would have left three.
5193 assert_eq!(
5194 lower(&mut names, &func),
5195 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
5196 %1:gpr($rsi) = x64.arg_val_64\n %2:gpr = x64.lea_64 [%0 + %1*4]\n \
5197 %3:gpr = x64.lea_64 [%2 + %1*4]\n}\n"
5198 );
5199 }
5200
5201 #[test]
5202 fn an_instruction_one_of_its_readers_cannot_take_is_folded_into_none_of_them() {
5203 let i64 = Type::int(64);
5204 let (mut names, mut func, block, args) = blank(&[i64, i64]);
5205 let mut build = Builder::new(&mut func, block);
5206 let four = build.iconst(i64, 4);
5207 let scaled = build.binary(Opcode::Mul, args[1], four, Flags::default());
5208 build.binary(Opcode::Add, args[0], scaled, Flags::default());
5209 build.store(scaled, args[0], plain(), Flags::default());
5210
5211 // The addition has room for the multiply and the store does not: what a store writes is
5212 // a register, and no rule reaches through it. Folding into the addition alone would
5213 // leave the multiply where it is for the store to read and do the work twice, so the
5214 // multiply is put back and both readers read the register it wrote.
5215 let text = lower(&mut names, &func);
5216 assert!(text.contains("x64.lea_64 [%1*4]"), "{text}");
5217 assert!(text.contains("x64.add_rr_64"), "{text}");
5218 }
5219
5220 #[test]
5221 fn a_shift_by_a_register_asks_for_it_in_cl() {
5222 let i32 = Type::int(32);
5223 let (mut names, mut func, block, args) = blank(&[i32, i32]);
5224 let mut build = Builder::new(&mut func, block);
5225 build.binary(Opcode::Shl, args[0], args[1], Flags::default());
5226
5227 // The fixed register is not in the rule. It is what the target says the instruction does
5228 // with its operands, and the allocator is what will act on it.
5229 let text = lower(&mut names, &func);
5230 assert!(text.contains("x64.shl_rcl_32 %0, %1($rcx)"), "{text}");
5231 }
5232
5233 #[test]
5234 fn a_division_names_the_registers_and_the_register_it_destroys() {
5235 let i32 = Type::int(32);
5236 let (mut names, mut func, block, args) = blank(&[i32, i32]);
5237 let mut build = Builder::new(&mut func, block);
5238 build.binary(Opcode::SDiv, args[0], args[1], Flags::default());
5239
5240 // Two definitions, because a division writes the remainder whether anybody wanted it or
5241 // not, and the second one is early because it is destroyed before the operands are read.
5242 let text = lower(&mut names, &func);
5243 assert!(
5244 text.contains("%2:gpr($rax), early %3:gpr($rdx) = x64.idiv_quo_32 %0($rax), %1"),
5245 "{text}"
5246 );
5247 }
5248
5249 #[test]
5250 fn a_load_reads_through_the_register_the_address_is_in() {
5251 let i64 = Type::int(64);
5252 let (mut names, mut func, block, args) = blank(&[i64]);
5253 let mut build = Builder::new(&mut func, block);
5254 build.load(Type::int(32), args[0], plain(), Flags::default());
5255
5256 assert_eq!(
5257 lower(&mut names, &func),
5258 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
5259 %1:gpr = x64.mov_rm_32 [%0]\n}\n"
5260 );
5261 }
5262
5263 #[test]
5264 fn a_store_writes_no_register_and_the_value_it_writes_is_the_one_the_ir_gave_it() {
5265 let (mut names, mut func, block, args) = blank(&[Type::int(32), Type::int(64)]);
5266 let mut build = Builder::new(&mut func, block);
5267 build.store(args[0], args[1], plain(), Flags::default());
5268
5269 // The value is the first parameter and the address is the second, and the instruction
5270 // takes them the other way round. Getting that backwards would compile to a store of the
5271 // address into the value, which is a program that runs and does the wrong thing.
5272 assert_eq!(
5273 lower(&mut names, &func),
5274 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
5275 %1:gpr($rsi) = x64.arg_val_64\n x64.mov_mr_32 %0, [%1]\n}\n"
5276 );
5277 }
5278
5279 #[test]
5280 fn an_address_with_a_constant_added_folds_into_the_access() {
5281 let i64 = Type::int(64);
5282 let (mut names, mut func, block, args) = blank(&[i64]);
5283 let mut build = Builder::new(&mut func, block);
5284 let twelve = build.iconst(i64, 12);
5285 let field = build.binary(Opcode::Add, args[0], twelve, Flags::default());
5286 build.load(Type::int(64), field, plain(), Flags::default());
5287
5288 // Two IR instructions and one machine instruction, which is what every read of a field
5289 // of a structure comes to.
5290 assert_eq!(
5291 lower(&mut names, &func),
5292 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
5293 %1:gpr = x64.mov_rm_64 [%0 + 12]\n}\n"
5294 );
5295 }
5296
5297 #[test]
5298 fn a_displacement_too_wide_to_encode_leaves_the_addition_where_it_is() {
5299 let i64 = Type::int(64);
5300 let (mut names, mut func, block, args) = blank(&[i64]);
5301 let mut build = Builder::new(&mut func, block);
5302 let big = build.iconst(i64, i128::from(i32::MAX) + 1);
5303 let far = build.binary(Opcode::Add, args[0], big, Flags::default());
5304 build.load(Type::int(32), far, plain(), Flags::default());
5305
5306 // A displacement is signed and 32 bits. The rule that folds one has a guard that turns
5307 // this down, so the addition stays and the load reads through what it produced. Nobody
5308 // wrote that fallback: it is the next way of showing the operand.
5309 let text = lower(&mut names, &func);
5310 assert!(text.contains("x64.mov_rm_32 [%2]"), "{text}");
5311 assert!(text.contains("x64.add_rr_64"), "{text}");
5312 }
5313
5314 #[test]
5315 fn a_store_of_a_value_that_was_loaded_is_two_instructions_and_no_arithmetic() {
5316 let i64 = Type::int(64);
5317 let (mut names, mut func, block, args) = blank(&[i64, i64]);
5318 let mut build = Builder::new(&mut func, block);
5319 let got = build.load(Type::int(8), args[0], plain(), Flags::default());
5320 build.store(got, args[1], plain(), Flags::default());
5321
5322 // A load feeding a store is the one place folding would be wrong: an x86-64 `mov` has at
5323 // most one memory operand, and there is no rule that takes two, so the load is left where
5324 // it is and the store reads the register it wrote.
5325 assert_eq!(
5326 lower(&mut names, &func),
5327 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
5328 %1:gpr($rsi) = x64.arg_val_64\n %2:gpr = x64.mov_rm_8 [%0]\n \
5329 x64.mov_mr_8 %2, [%1]\n}\n"
5330 );
5331 }
5332
5333 #[test]
5334 fn an_access_at_a_width_no_rule_is_written_at_is_reported() {
5335 let i64 = Type::int(64);
5336 let (mut names, mut source, block, args) = blank(&[i64]);
5337 let mut build = Builder::new(&mut source, block);
5338 build.load(Type::int(128), args[0], plain(), Flags::default());
5339
5340 // The width is the whole of what is wrong here, so the width is in the message: `load`
5341 // on its own is written about at every other width and would send a reader looking in
5342 // the wrong place.
5343 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
5344 .expect_err("nothing loads 128 bits");
5345 assert_eq!(failed.to_string(), "no rule lowers a `load` producing a `i128`");
5346 }
5347
5348 #[test]
5349 fn a_return_asks_for_the_value_in_the_register_the_caller_reads() {
5350 let (mut names, mut func, block, args) = blank(&[Type::int(32)]);
5351 let mut build = Builder::new(&mut func, block);
5352 build.ret(&[args[0]]);
5353
5354 // The register is not in the rule, the same way `cl` is not in the rule for a shift. It
5355 // is what the target says the instruction does with its operand, and the allocator is
5356 // what will act on it. There is no `ret` here, because giving the frame back has to
5357 // happen between this and leaving and the frame is not worked out yet.
5358 assert_eq!(
5359 lower(&mut names, &func),
5360 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
5361 x64.ret_val_32 %0($rax)\n}\n"
5362 );
5363 }
5364
5365 #[test]
5366 fn a_return_of_two_values_asks_for_the_second_register_as_well() {
5367 let i64 = Type::int(64);
5368 let (mut names, mut func, block, args) = blank(&[i64, i64]);
5369 let mut build = Builder::new(&mut func, block);
5370 build.ret(&[args[0], args[1]]);
5371
5372 // `struct { long a, b; } f(long a, long b)`, after the front end has classified it. Both
5373 // halves are integers, so the second is in the second integer return register, and both
5374 // pseudos say so the same way the one for a single value does.
5375 assert_eq!(
5376 lower(&mut names, &func),
5377 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
5378 %1:gpr($rsi) = x64.arg_val_64\n x64.ret_val_64 %0($rax)\n \
5379 x64.ret_val2_64 %1($rdx)\n}\n"
5380 );
5381 }
5382
5383 #[test]
5384 fn two_values_back_in_different_files_are_both_the_first_of_their_own() {
5385 let f64 = Type::float(rucc_ir::Float::F64);
5386 let (mut names, mut func, block, args) = blank(&[f64, Type::int(64)]);
5387 let mut build = Builder::new(&mut func, block);
5388 build.ret(&[args[0], args[1]]);
5389
5390 // `struct { double a; long b; } f(double a, long b)`. The two files are counted apart, so
5391 // neither half is the second of anything and the `double` is in `xmm0` rather than in the
5392 // register a second `double` would have been in. Getting this wrong is not a crash: the
5393 // caller reads a register nobody wrote, and this is where that is ruled out.
5394 assert_eq!(
5395 lower(&mut names, &func),
5396 "mfunc @f {\nblock0:\n %0:xmm($xmm0) = x64.arg_val_f64\n \
5397 %1:gpr($rdi) = x64.arg_val_64\n x64.ret_val_f64 %0($xmm0)\n \
5398 x64.ret_val_64 %1($rax)\n}\n"
5399 );
5400 }
5401
5402 #[test]
5403 fn two_of_the_same_file_back_take_the_first_two_of_it() {
5404 let f64 = Type::float(rucc_ir::Float::F64);
5405 let (mut names, mut func, block, args) = blank(&[f64, f64]);
5406 let mut build = Builder::new(&mut func, block);
5407 build.ret(&[args[0], args[1]]);
5408
5409 // `struct { double x, y; } f(double x, double y)`, which is the vector half of the pair
5410 // above and counts in its own file the same way.
5411 assert_eq!(
5412 lower(&mut names, &func),
5413 "mfunc @f {\nblock0:\n %0:xmm($xmm0) = x64.arg_val_f64\n \
5414 %1:xmm($xmm1) = x64.arg_val_f64\n x64.ret_val_f64 %0($xmm0)\n \
5415 x64.ret_val2_f64 %1($xmm1)\n}\n"
5416 );
5417 }
5418
5419 /// A function whose answer goes back through memory, with the pointer to the space for it in
5420 /// front of whatever else it takes. Only the signature says it is one.
5421 fn returning_through_memory(params: &[Type]) -> (Interner, Func, Block, Vec<Value>) {
5422 let mut names = Interner::new();
5423 let sret = Abi::Sret { size: 32, align: 8 };
5424 let mut signature = Signature::new().and_param(Param::with_abi(Type::PTR, sret));
5425 signature.params.extend(params.iter().copied().map(Param::new));
5426 let mut func = Func::new(names.intern("f"), signature);
5427 let block = func.create_block();
5428 let space = func.append_param(block, Type::PTR);
5429 let values = std::iter::once(space)
5430 .chain(params.iter().map(|&ty| func.append_param(block, ty)))
5431 .collect();
5432 (names, func, block, values)
5433 }
5434
5435 #[test]
5436 fn the_space_a_return_through_memory_was_given_goes_back_in_the_first_return_register() {
5437 let (mut names, mut func, block, _) = returning_through_memory(&[]);
5438 Builder::new(&mut func, block).ret(&[]);
5439
5440 // `struct big f(void)`, where `big` is too large to come back in registers. The `return`
5441 // carries nothing, because the value went into the space the caller handed over, and the
5442 // document still says that address comes back in `rax`. Nothing in the IR says it, so the
5443 // convention says it, and the pseudo is the one any other pointer return would use.
5444 assert_eq!(
5445 lower(&mut names, &func),
5446 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
5447 x64.ret_val_64 %0($rax)\n}\n"
5448 );
5449 }
5450
5451 #[test]
5452 fn what_the_function_did_in_between_does_not_take_the_register_off_it() {
5453 let (mut names, mut func, block, args) = returning_through_memory(&[Type::int(32)]);
5454 let mut build = Builder::new(&mut func, block);
5455 build.store(args[1], args[0], plain(), Flags::default());
5456 build.ret(&[]);
5457
5458 // The register is a read at the end and not a move at the start, so it is live across
5459 // everything between the two and the allocator has to keep it somewhere. In a function
5460 // with a call in it that somewhere is a callee saved register, and the address comes back
5461 // into `rax` here rather than whatever the last instruction happened to leave there. That
5462 // is issue #333, and a store is enough to show the value outlives the entry block.
5463 let text = lower(&mut names, &func);
5464 assert!(text.contains("x64.mov_mr_32 %1, [%0]"), "{text}");
5465 assert!(text.ends_with(" x64.ret_val_64 %0($rax)\n}\n"), "{text}");
5466 }
5467
5468 #[test]
5469 fn a_pointer_that_is_only_a_pointer_is_not_given_back() {
5470 let (mut names, mut func, block, args) = blank(&[Type::PTR]);
5471 let mut build = Builder::new(&mut func, block);
5472 build.store(args[0], args[0], plain(), Flags::default());
5473 build.ret(&[]);
5474
5475 // `void f(void **p)`. It takes a pointer first and returns nothing, which is the shape of
5476 // the one above and none of its meaning, and what tells them apart is the signature. A
5477 // `void` function leaves `rax` alone.
5478 assert!(!lower(&mut names, &func).contains("ret_val"));
5479 }
5480
5481 #[test]
5482 fn a_return_of_a_constant_puts_it_in_a_register_first() {
5483 let (mut names, mut func, block, _) = blank(&[]);
5484 let mut build = Builder::new(&mut func, block);
5485 let zero = build.iconst(Type::int(32), 0);
5486 build.ret(&[zero]);
5487
5488 // No rule returns an immediate, so the plan that offers one is turned down and the next
5489 // one materializes it. That is `int main(void) { return 0; }` in full, once the epilogue
5490 // is appended to it.
5491 assert_eq!(
5492 lower(&mut names, &func),
5493 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_ri_32 0\n x64.ret_val_32 %0($rax)\n}\n"
5494 );
5495 }
5496
5497 #[test]
5498 fn the_rule_that_writes_a_constant_down_is_recorded_as_a_rule_that_fired() {
5499 let (mut names, mut func, block, _) = blank(&[]);
5500 let mut build = Builder::new(&mut func, block);
5501 let zero = build.iconst(Type::int(32), 0);
5502 build.ret(&[zero]);
5503
5504 // The loop over the instructions passes a constant by, because a constant is written where
5505 // a register for it is first wanted rather than where the IR put it. So the only place a
5506 // rule about one is ever selected is the materialization, and a mark made in the loop
5507 // alone would report every rule about a constant as a rule nothing reaches.
5508 let out = super::func(&func, &mut names, &SYSV, &Elsewhere::default())
5509 .expect("every instruction has a rule");
5510 let rules = &crate::select::x86_64::TABLE.rules;
5511 let fired: Vec<&str> = rules
5512 .iter()
5513 .enumerate()
5514 .filter(|(index, _)| out.fired.has(*index))
5515 .map(|(_, rule)| rule.pattern)
5516 .collect();
5517 assert!(fired.contains(&"(iconst.i32 k)"), "{fired:?}");
5518 }
5519
5520 #[test]
5521 fn a_return_of_nothing_is_no_instruction_at_all() {
5522 let (mut names, mut func, block, _) = blank(&[]);
5523 let mut build = Builder::new(&mut func, block);
5524 build.ret(&[]);
5525
5526 // Every part of leaving a function that returns nothing is the epilogue's, and the
5527 // epilogue goes in after allocation. A block with nothing in it is the right answer here
5528 // rather than a function that could not be lowered.
5529 assert_eq!(lower(&mut names, &func), "mfunc @f {\nblock0:\n}\n");
5530 }
5531
5532 #[test]
5533 fn the_allocator_is_what_moves_the_answer_into_the_return_register() {
5534 let (mut names, mut source, block, _) = blank(&[]);
5535 let mut build = Builder::new(&mut source, block);
5536 let zero = build.iconst(Type::int(32), 0);
5537 build.ret(&[zero]);
5538
5539 let mut out = func(&source, &mut names, &SYSV, &Elsewhere::default())
5540 .expect("every instruction has a rule")
5541 .func;
5542 let env = env();
5543 let allocation = rucc_regalloc::run(&mut out, &env, "test", true);
5544 let frame = Frame::of(&out, &allocation, &Layout::new(&SYSV, REGS));
5545 finish(
5546 &mut out,
5547 &allocation,
5548 &frame,
5549 &Stack::default(),
5550 Convention::new(&SYSV, &FRAME),
5551 &mut names,
5552 );
5553
5554 // `int main(void) { return 0; }` end to end. Nothing here asked for `rax`: the rule said
5555 // the value goes back, the target said where, and the allocator is what made it true. The
5556 // epilogue is what leaves, and this function needs no frame, so it is the return alone.
5557 //
5558 // Two instructions and no copy, which is what a hint buys. The return insists on `rax`,
5559 // so `rax` is the register the allocator tries first for the value the return reads, and
5560 // the constant is written straight into it.
5561 assert_eq!(
5562 mir::print_func(&out, &names, ®S),
5563 "mfunc @f {\nblock0:\n $rax = x64.mov_ri_32 0\n \
5564 x64.ret_val_32 $rax($rax)\n x64.ret\n}\n"
5565 );
5566 }
5567
5568 #[test]
5569 fn a_function_of_two_arguments_is_a_whole_function_now() {
5570 let i32 = Type::int(32);
5571 let (mut names, mut source, block, args) = blank(&[i32, i32]);
5572 let mut build = Builder::new(&mut source, block);
5573 let sum = build.binary(Opcode::Add, args[0], args[1], Flags::default());
5574 build.ret(&[sum]);
5575
5576 let mut out = func(&source, &mut names, &SYSV, &Elsewhere::default())
5577 .expect("every instruction has a rule")
5578 .func;
5579 let env = env();
5580 let allocation = rucc_regalloc::run(&mut out, &env, "test", true);
5581 let frame = Frame::of(&out, &allocation, &Layout::new(&SYSV, REGS));
5582 finish(
5583 &mut out,
5584 &allocation,
5585 &frame,
5586 &Stack::default(),
5587 Convention::new(&SYSV, &FRAME),
5588 &mut names,
5589 );
5590
5591 // `int f(int a, int b) { return a + b; }` end to end, and this is the test the argument
5592 // side exists for. Before it there was no way to write one: the allocator refuses a
5593 // function whose entry block takes parameters, because there is no edge into an entry
5594 // block for the moves that give a block parameter its value to go on.
5595 //
5596 // One move, and it is the one the machine's addition needs rather than one the allocator
5597 // owes anybody. Each argument stays in the register it arrived in, because the pseudo
5598 // that defines it insists on that register and the allocator now tries it first, and the
5599 // sum stays in the register the addition wrote it to until the return reads it out. The
5600 // copy in front of a two address instruction is what makes its destination one of the
5601 // registers it reads, and the source operand keeps its own name because the destination
5602 // is what the encoder writes.
5603 assert_eq!(
5604 mir::print_func(&out, &names, ®S),
5605 "mfunc @f {\nblock0:\n $rdi($rdi) = x64.arg_val_32\n \
5606 $rsi($rsi) = x64.arg_val_32\n \
5607 $rdi(reuse 1) = x64.add_rr_32 $rdi, $rsi\n $rax = x64.mov_rr_64 $rdi\n \
5608 x64.ret_val_32 $rax($rax)\n x64.ret\n}\n"
5609 );
5610 }
5611
5612 #[test]
5613 fn an_argument_with_no_register_left_for_it_is_read_out_of_the_caller_s_stack() {
5614 let i64 = Type::int(64);
5615 let (mut names, mut source, block, args) = blank(&[i64; 7]);
5616 let mut build = Builder::new(&mut source, block);
5617 build.ret(&[args[6]]);
5618
5619 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
5620 .expect("the seventh is read from memory");
5621
5622 // SysV passes six integers in registers and the seventh in the caller's memory, so six of
5623 // these are pseudos that encode to nothing and the seventh is a load that encodes to real
5624 // bytes. Its displacement is nothing here for the reason a local's is: there is no frame
5625 // yet. What the walk hands on is which instruction is waiting, and for how far up the
5626 // caller's argument area, which is the bottom of it because it is the first one there.
5627 assert_eq!(lowered.stack.arguments.len(), 1);
5628 assert_eq!(lowered.stack.arguments[0].1, 0);
5629 let text = mir::print_func(&lowered.func, &names, ®S);
5630 assert!(text.contains("%6:gpr = x64.mov_rm_64 [$rsp]"), "{text}");
5631 assert_eq!(text.matches("x64.arg_val_64").count(), 6, "{text}");
5632 }
5633
5634 #[test]
5635 fn the_frame_is_what_says_how_far_up_the_caller_s_stack_an_argument_is() {
5636 let i64 = Type::int(64);
5637 let (mut names, mut source, block, args) = blank(&[i64; 8]);
5638 let mut build = Builder::new(&mut source, block);
5639 let sum = build.binary(Opcode::Add, args[6], args[7], Flags::default());
5640 build.ret(&[sum]);
5641
5642 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
5643 .expect("both are read from memory");
5644 let stack = lowered.stack;
5645 let mut out = lowered.func;
5646 let env = env();
5647 let allocation = rucc_regalloc::run(&mut out, &env, "test", true);
5648 let layout = stack.layout(Layout::new(&SYSV, REGS));
5649 let frame = Frame::of(&out, &allocation, &layout);
5650 finish(&mut out, &allocation, &frame, &stack, Convention::new(&SYSV, &FRAME), &mut names);
5651
5652 // A leaf that takes no frame, so the stack pointer never moves and the only thing between
5653 // it and the caller's arguments is the return address the call pushed. The seventh
5654 // parameter is at the bottom of the caller's argument area and the eighth is one word
5655 // further up, which is the eight bytes between the two offsets.
5656 let text = mir::print_func(&out, &names, ®S);
5657 assert_eq!(frame.size(), 0);
5658 assert_eq!(frame.incoming(), Incoming::from_stack(8));
5659 assert!(text.contains("x64.mov_rm_64 [$rsp + 8]"), "{text}");
5660 assert!(text.contains("x64.mov_rm_64 [$rsp + 16]"), "{text}");
5661 }
5662
5663 #[test]
5664 fn a_realigned_frame_reaches_the_caller_s_arguments_through_the_frame_pointer() {
5665 let i64 = Type::int(64);
5666 let (mut names, mut source, block, args) = blank(&[i64; 7]);
5667 let wide = slot(&mut source, block, 64, 32);
5668 let mut build = Builder::new(&mut source, block);
5669 build.store(args[6], wide, plain(), Flags::default());
5670 build.ret(&[args[6]]);
5671
5672 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
5673 .expect("every instruction has a rule");
5674 let stack = lowered.stack;
5675 let mut out = lowered.func;
5676 let env = env();
5677 let allocation = rucc_regalloc::run(&mut out, &env, "test", true);
5678 let layout = stack.layout(Layout::new(&SYSV, REGS));
5679 let frame = Frame::of(&out, &allocation, &layout);
5680 finish(&mut out, &allocation, &frame, &stack, Convention::new(&SYSV, &FRAME), &mut names);
5681
5682 // A local wanting thirty two byte alignment makes the prologue force the stack pointer,
5683 // which throws away how far the caller's stack was. So the load the lowering wrote off the
5684 // stack pointer is rewritten to read through the frame pointer, at the one distance that
5685 // survives: the word the prologue pushed the frame pointer into, and the return address
5686 // above it.
5687 let text = mir::print_func(&out, &names, ®S);
5688 assert_eq!(frame.realign(), Some(32));
5689 assert_eq!(frame.incoming(), Incoming::from_frame(16));
5690 assert!(text.contains("x64.mov_rm_64 [$rbp + 16]"), "{text}");
5691 assert!(!text.contains("x64.mov_rm_64 [$rsp"), "{text}");
5692 }
5693
5694 #[test]
5695 fn a_jump_is_the_edge_and_nothing_else() {
5696 let i32 = Type::int(32);
5697 let (mut names, mut source, entry, args) = blank(&[i32]);
5698 let next = source.create_block();
5699 let got = source.append_param(next, i32);
5700 Builder::new(&mut source, entry).jump(next, &[args[0]]);
5701 Builder::new(&mut source, next).ret(&[got]);
5702
5703 // Two blocks and two instructions, and the jump is neither of them. What it was is the
5704 // arm on the first block, and what the arm carries is the argument it was called with.
5705 assert_eq!(
5706 lower(&mut names, &source),
5707 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32 block1(%0)\n\n\
5708 block1(%1:gpr):\n x64.ret_val_32 %1($rax)\n}\n"
5709 );
5710 }
5711
5712 /// A block that reads what a block below it writes is filled after it, not before it.
5713 ///
5714 /// The blocks are written entry, `early`, `late`, `exit`, and the entry jumps straight past
5715 /// `early` to `late`, so `late` dominates `early` while sitting below it in the function.
5716 /// Filling them in the order they are written reaches the read in `early` first, and reading
5717 /// a value with no register yet mints one. The cast in `late` is no instruction at all, so
5718 /// what it does is give its answer the register its operand is already in, and that is not
5719 /// the register the read minted. Nothing writes the register the read minted. The printer
5720 /// says `%?` for a register nothing defines, which is what this looks for, and what came out
5721 /// of the real bug was SQLite loading a stack slot no store ever reached.
5722 #[test]
5723 fn a_block_that_reads_what_a_block_below_it_writes_is_filled_after_it() {
5724 let i64 = Type::int(64);
5725 let (mut names, mut source, entry, args) = blank(&[i64, i64]);
5726 let early = source.create_block();
5727 let late = source.create_block();
5728 let exit = source.create_block();
5729
5730 Builder::new(&mut source, entry).jump(late, &[]);
5731 let ptr = cast(&mut source, late, Opcode::IntToPtr, args[0], Type::PTR);
5732 Builder::new(&mut source, early).ret(&[ptr]);
5733 let mut build = Builder::new(&mut source, late);
5734 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
5735 build.br_if(cond, early, &[], exit, &[]);
5736 Builder::new(&mut source, exit).ret(&[args[1]]);
5737
5738 let text = lower(&mut names, &source);
5739 assert!(!text.contains("%?"), "every register has something that writes it: {text}");
5740 }
5741
5742 /// A constant is written where it is wanted rather than where the IR defined it, and two
5743 /// blocks wanting the same one is two places. Writing it once and reading it in both is a
5744 /// register read where nothing wrote it, unless the block it was written in happens to
5745 /// dominate the other, which nothing here checks and which the second arm of a branch never
5746 /// does. Each block gets its own copy of the number instead.
5747 #[test]
5748 fn a_constant_two_blocks_want_is_written_in_both_of_them() {
5749 let i32 = Type::int(32);
5750 let (mut names, mut source, entry, args) = blank(&[i32, i32]);
5751 let then = source.create_block();
5752 let other = source.create_block();
5753 let join = source.create_block();
5754 let got = source.append_param(join, i32);
5755
5756 let mut build = Builder::new(&mut source, entry);
5757 let seven = build.iconst(i32, 7);
5758 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
5759 build.br_if(cond, then, &[], other, &[]);
5760 // Both arms want the seven in a register, because a block argument is never an immediate,
5761 // and neither arm dominates the other.
5762 Builder::new(&mut source, then).jump(join, &[seven]);
5763 Builder::new(&mut source, other).jump(join, &[seven]);
5764 Builder::new(&mut source, join).ret(&[got]);
5765
5766 let text = lower(&mut names, &source);
5767 assert_eq!(text.matches("x64.mov_ri_32 7").count(), 2, "one seven per block: {text}");
5768 }
5769
5770 /// An argument on an edge out of a block that leaves two ways is read after every instruction
5771 /// of the block is written, and reading one can write an instruction, which would land after
5772 /// the branch that has already jumped past it. The branch goes back on the end.
5773 #[test]
5774 fn a_constant_an_edge_wants_is_written_before_the_branch_and_not_after_it() {
5775 let i32 = Type::int(32);
5776 let (mut names, mut source, entry, args) = blank(&[i32, i32]);
5777 let then = source.create_block();
5778 let join = source.create_block();
5779 let got = source.append_param(join, i32);
5780
5781 let mut build = Builder::new(&mut source, entry);
5782 let nine = build.iconst(i32, 9);
5783 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
5784 build.br_if(cond, then, &[], join, &[nine]);
5785 Builder::new(&mut source, then).jump(join, &[args[0]]);
5786 Builder::new(&mut source, join).ret(&[got]);
5787
5788 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
5789 .expect("every instruction has a rule")
5790 .func;
5791 let entry = out.entry().expect("an entry block");
5792 let last = out.terminator(entry).expect("a block that leaves two ways has a branch");
5793 let branch = names.intern("x64.br_cond_8");
5794 assert_eq!(
5795 out[last].opcode,
5796 mir::Opcode::new(branch),
5797 "the branch is last: {}",
5798 mir::print_func(&out, &names, ®S)
5799 );
5800 }
5801
5802 #[test]
5803 fn a_conditional_branch_is_lowered_to_the_condition_and_nothing_about_where_it_goes() {
5804 let i32 = Type::int(32);
5805 let (mut names, mut source, entry, args) = blank(&[i32, i32]);
5806 let then = source.create_block();
5807 let other = source.create_block();
5808 let mut build = Builder::new(&mut source, entry);
5809 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
5810 build.br_if(cond, then, &[], other, &[]);
5811 Builder::new(&mut source, then).ret(&[args[0]]);
5812 Builder::new(&mut source, other).ret(&[args[1]]);
5813
5814 // The comparison writes a byte and the branch reads it, and neither says a block. Both
5815 // arms are on the entry block, in the order the branch took them, so the arm that runs
5816 // when the condition holds is the first.
5817 assert_eq!(
5818 lower(&mut names, &source),
5819 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
5820 %1:gpr($rsi) = x64.arg_val_32\n %2:gpr = x64.cmp_set_l_32 %0, %1\n \
5821 x64.br_cond_8 %2, block1, block2\n\n\
5822 block1:\n x64.ret_val_32 %0($rax)\n\n\
5823 block2:\n x64.ret_val_32 %1($rax)\n}\n"
5824 );
5825 }
5826
5827 /// A choice between two values, which is one instruction and no blocks at all.
5828 ///
5829 /// The arms come out the other way round from the IR, because a conditional move overwrites its
5830 /// destination and the destination is the arm taken when the condition does not hold. The
5831 /// condition arrives last for the same reason: it is read by the test in front of the move
5832 /// rather than by the move.
5833 #[test]
5834 fn a_select_is_lowered_to_a_test_and_a_conditional_move() {
5835 let i32 = Type::int(32);
5836 let (mut names, mut source, entry, args) = blank(&[i32, i32]);
5837 let mut build = Builder::new(&mut source, entry);
5838 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
5839 let picked = build.select(cond, args[0], args[1]);
5840 build.ret(&[picked]);
5841
5842 assert_eq!(
5843 lower(&mut names, &source),
5844 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
5845 %1:gpr($rsi) = x64.arg_val_32\n %2:gpr = x64.cmp_set_l_32 %0, %1\n \
5846 %3:gpr(reuse 1) = x64.test_cmov_ne_32 %1, %0, %2\n \
5847 x64.ret_val_32 %3($rax)\n}\n"
5848 );
5849 }
5850
5851 #[test]
5852 fn a_branch_over_a_block_is_a_whole_function_now() {
5853 let i32 = Type::int(32);
5854 let (mut names, mut source, entry, args) = blank(&[i32, i32]);
5855 let then = source.create_block();
5856 let other = source.create_block();
5857 let join = source.create_block();
5858 let got = source.append_param(join, i32);
5859 let mut build = Builder::new(&mut source, entry);
5860 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
5861 build.br_if(cond, then, &[], other, &[]);
5862 let mut build = Builder::new(&mut source, then);
5863 let sum = build.binary(Opcode::Add, args[0], args[1], Flags::default());
5864 build.jump(join, &[sum]);
5865 Builder::new(&mut source, other).jump(join, &[args[1]]);
5866 Builder::new(&mut source, join).ret(&[got]);
5867
5868 // `int f(int a, int b) { if (a < b) return a + b; else return b; }` end to end, written
5869 // the way a front end writes it: both arms of the branch are blocks of their own and the
5870 // return is the block they meet at. No edge here is critical, because the two arms out of
5871 // the entry carry nothing and the two arms into the join each leave a block that goes
5872 // nowhere else, so each has its own end to put its move at.
5873 let mut out = func(&source, &mut names, &SYSV, &Elsewhere::default())
5874 .expect("every instruction has a rule")
5875 .func;
5876 assert_eq!(crate::split::critical(&mut out), 0, "no edge here is critical");
5877 let env = env();
5878 let allocation = rucc_regalloc::run(&mut out, &env, "test", true);
5879 let frame = Frame::of(&out, &allocation, &Layout::new(&SYSV, REGS));
5880 finish(
5881 &mut out,
5882 &allocation,
5883 &frame,
5884 &Stack::default(),
5885 Convention::new(&SYSV, &FRAME),
5886 &mut names,
5887 );
5888
5889 // One epilogue, on the join, which is the one block the function leaves from, and the
5890 // moves that give the join its parameter are at the end of each arm. Every register is
5891 // physical and the branch is still a branch on a register, because turning it into a
5892 // `test` and a `jcc` is the block layout's and there is no block layout yet.
5893 let text = mir::print_func(&out, &names, ®S);
5894 assert_eq!(text.matches("x64.ret\n").count(), 1, "{text}");
5895 assert!(text.contains("x64.br_cond_8"), "{text}");
5896 assert!(text.contains("x64.add_rr_32"), "{text}");
5897 assert!(!text.contains('%'), "{text}");
5898 }
5899
5900 #[test]
5901 fn a_critical_edge_is_split_before_the_allocator_ever_sees_it() {
5902 let i32 = Type::int(32);
5903 let (mut names, mut source, entry, args) = blank(&[i32, i32]);
5904 let then = source.create_block();
5905 let join = source.create_block();
5906 let got = source.append_param(join, i32);
5907 let mut build = Builder::new(&mut source, entry);
5908 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
5909 build.br_if(cond, then, &[], join, &[args[1]]);
5910 Builder::new(&mut source, then).jump(join, &[args[0]]);
5911 let mut build = Builder::new(&mut source, join);
5912 let twice = build.binary(Opcode::Add, got, got, Flags::default());
5913 build.ret(&[twice]);
5914
5915 // The else arm is critical: the entry block leaves two ways and the join is arrived at
5916 // two ways, and the arm carries a value. Without splitting it the allocator asserts,
5917 // because the move that gives the join its parameter would have to run at the end of a
5918 // block that also goes to the other arm.
5919 let mut out = func(&source, &mut names, &SYSV, &Elsewhere::default())
5920 .expect("every instruction has a rule")
5921 .func;
5922 assert_eq!(crate::split::critical(&mut out), 1);
5923 let env = env();
5924 let allocation = rucc_regalloc::run(&mut out, &env, "test", true);
5925 let frame = Frame::of(&out, &allocation, &Layout::new(&SYSV, REGS));
5926 finish(
5927 &mut out,
5928 &allocation,
5929 &frame,
5930 &Stack::default(),
5931 Convention::new(&SYSV, &FRAME),
5932 &mut names,
5933 );
5934
5935 // The block the split added is where the move went, and it is the whole of that block.
5936 let text = mir::print_func(&out, &names, ®S);
5937 assert_eq!(out.block_count(), 4, "{text}");
5938 assert_eq!(text.matches("x64.ret\n").count(), 1, "{text}");
5939 }
5940
5941 #[test]
5942 fn a_call_passes_what_the_convention_says_and_takes_back_what_it_says() {
5943 let i32 = Type::int(32);
5944 let (mut names, mut source, block, args) = blank(&[i32, i32]);
5945 let sig =
5946 source.add_signature(Signature::new().with_params(&[i32, i32]).with_returns(&[i32]));
5947 let callee = names.intern("g");
5948 let call = Builder::new(&mut source, block).call(callee, sig, &[args[0], args[1]]);
5949 let got = source[call].first_result.expect("an integer comes back");
5950 Builder::new(&mut source, block).ret(&[got]);
5951
5952 // `int f(int a, int b) { return g(a, b); }`. The arguments arrived where the call wants
5953 // them, so what the call reads is what arrived, and the whole of the convention is in the
5954 // constraints rather than in a move.
5955 let text = lower(&mut names, &source);
5956 assert!(text.contains("= x64.call %0($rdi), %1($rsi), @g"), "{text}");
5957 assert!(text.contains("x64.ret_val_32 %2($rax)"), "{text}");
5958 // What the call writes is the value that comes back and then every register the callee is
5959 // free to destroy, in both classes, which is the whole of what stops the allocator from
5960 // leaving something in one of them.
5961 assert!(text.contains("%2:gpr($rax), $rcx, $rdx, $r8, $r9, $r10, $r11, $xmm0,"), "{text}");
5962 assert!(text.contains("$xmm15 = x64.call"), "{text}");
5963 }
5964
5965 #[test]
5966 fn what_the_frame_owes_a_call_comes_back_with_the_function() {
5967 let i32 = Type::int(32);
5968 let sig = |source: &mut Func| source.add_signature(Signature::new().with_params(&[i32]));
5969
5970 let (mut names, mut source, block, args) = blank(&[i32]);
5971 let sig = sig(&mut source);
5972 let callee = names.intern("g");
5973 Builder::new(&mut source, block).call(callee, sig, &[args[0]]);
5974 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
5975 .expect("every instruction has a rule");
5976
5977 // Nothing on the stack, so nothing owed, but not a leaf either: a function that calls
5978 // owes the callee an aligned stack pointer and may not use the red zone.
5979 assert_eq!(out.stack.calls, Some(0));
5980 let layout = out.stack.layout(Layout::new(&SYSV, REGS));
5981 assert!(!layout.leaf);
5982 assert_eq!(layout.outgoing, 0);
5983
5984 // The same call under the other convention owes thirty two bytes for the callee to spill
5985 // its register arguments into, which is a fact about the convention and not about the call.
5986 let out = func(&source, &mut names, &x86_64::WIN64, &Elsewhere::default())
5987 .expect("every instruction has a rule");
5988 assert_eq!(out.stack.calls, Some(32));
5989
5990 // And a function that calls nothing is a leaf, which is what says it may use the red zone.
5991 let (mut names, mut source, block, args) = blank(&[i32]);
5992 Builder::new(&mut source, block).ret(&[args[0]]);
5993 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
5994 .expect("every instruction has a rule");
5995 assert_eq!(out.stack.calls, None);
5996 assert!(out.stack.layout(Layout::new(&SYSV, REGS)).leaf);
5997 }
5998
5999 /// A Windows variadic prologue writes the argument registers the signature did not name into
6000 /// the shadow space the caller already reserved, which makes every argument one run of words up
6001 /// there and a `va_start` the address of the first of them. One `lea` and one store, and no
6002 /// counts, because a list that is a pointer has nowhere to put one and nothing that reads one.
6003 #[test]
6004 fn a_windows_variadic_function_homes_its_spare_registers_in_the_callers_area() {
6005 let mut names = Interner::new();
6006 let params = [Type::int(32), Type::PTR];
6007 let signature = Signature::new().with_params(¶ms).variadic();
6008 let mut source = Func::new(names.intern("f"), signature);
6009 let block = source.create_block();
6010 let values: Vec<Value> = params.iter().map(|&ty| source.append_param(block, ty)).collect();
6011 let mut build = Builder::new(&mut source, block);
6012 let args = build.func().push_values(&values[1..]);
6013 build.inst(InstData { args, ..InstData::new(Opcode::VaStart) }, &[]);
6014 build.ret(&[]);
6015
6016 let out = func(&source, &mut names, &x86_64::WIN64, &Elsewhere::default())
6017 .expect("every instruction has a rule");
6018 let text = mir::print_func(&out.func, &names, ®S);
6019
6020 // Two named parameters, so the registers at the next two positions hold arguments nobody
6021 // named and both are written up into the caller's area. The displacement is empty here and
6022 // `finish` fills it in, the same way it does for a parameter the registers ran out before.
6023 assert!(text.contains("($r8) = x64.arg_val_64"), "{text}");
6024 assert!(text.contains("($r9) = x64.arg_val_64"), "{text}");
6025 assert_eq!(text.matches("x64.mov_mr_64").count(), 3, "two homed and one stored: {text}");
6026 assert!(!text.contains("x64.mov_ri_32"), "and no field holds a count: {text}");
6027
6028 // All three waiting on the same fixup, and the last of them is the `lea` the list is given,
6029 // sixteen bytes up, which is where the two arguments the signature does name stopped.
6030 assert_eq!(out.stack.arguments.len(), 3);
6031 assert_eq!(out.stack.arguments[2].1, 16);
6032 }
6033
6034 #[test]
6035 fn a_value_that_outlives_a_call_is_not_left_where_the_call_destroys_it() {
6036 let i32 = Type::int(32);
6037 let (mut names, mut source, block, args) = blank(&[i32]);
6038 let sig = source.add_signature(Signature::new().with_params(&[i32]).with_returns(&[i32]));
6039 let callee = names.intern("g");
6040 let call = Builder::new(&mut source, block).call(callee, sig, &[args[0]]);
6041 let got = source[call].first_result.expect("an integer comes back");
6042 let mut build = Builder::new(&mut source, block);
6043 let sum = build.binary(Opcode::Add, got, args[0], Flags::default());
6044 build.ret(&[sum]);
6045
6046 // `int f(int a) { return g(a) + a; }`, which is the smallest program that asks the
6047 // question: `a` is read after the call and `rdi` is a register the call destroys.
6048 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
6049 .expect("every instruction has a rule");
6050 let layout = lowered.stack.layout(Layout::new(&SYSV, REGS));
6051 let mut out = lowered.func;
6052 let env = env();
6053 let allocation = rucc_regalloc::run(&mut out, &env, "test", true);
6054 let frame = Frame::of(&out, &allocation, &layout);
6055 finish(
6056 &mut out,
6057 &allocation,
6058 &frame,
6059 &Stack::default(),
6060 Convention::new(&SYSV, &FRAME),
6061 &mut names,
6062 );
6063
6064 // It went to a register the callee has to put back, and the prologue and epilogue are what
6065 // put it back, which is the whole bargain the two halves of a convention make.
6066 let text = mir::print_func(&out, &names, ®S);
6067 assert!(text.contains("$rbx"), "{text}");
6068 assert!(!text.contains('%'), "{text}");
6069 assert_eq!(text.matches("x64.call").count(), 1, "{text}");
6070 }
6071
6072 #[test]
6073 fn a_call_with_more_arguments_than_registers_writes_the_rest_into_the_outgoing_area() {
6074 let i64 = Type::int(64);
6075 let (mut names, mut source, block, args) = blank(&[i64]);
6076 let seven = vec![i64; 7];
6077 let sig = source.add_signature(Signature::new().with_params(&seven));
6078 let callee = names.intern("g");
6079 let passed = vec![args[0]; 7];
6080 Builder::new(&mut source, block).call(callee, sig, &passed);
6081
6082 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
6083 .expect("the seventh goes to memory");
6084 // The bytes the call needs are on the layout the frame is worked out from, so that the
6085 // frame reserves as many as the widest call in the function asked for.
6086 assert_eq!(lowered.stack.calls, Some(8));
6087 let text = mir::print_func(&lowered.func, &names, ®S);
6088 assert!(text.contains("x64.mov_mr_64 %0, [$rsp]\n"), "{text}");
6089 }
6090
6091 #[test]
6092 fn a_call_this_cannot_make_is_reported_rather_than_made() {
6093 let (mut names, mut source, block, _) = blank(&[]);
6094 let returns = [Type::float(rucc_ir::Float::F80), Type::int(64)];
6095 let sig = source.add_signature(Signature::new().with_returns(&returns));
6096 let callee = names.intern("g");
6097 Builder::new(&mut source, block).call(callee, sig, &[]);
6098 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
6099 .expect_err("a long double is on the x87");
6100 assert_eq!(failed.to_string(), "what this call gives back is on the x87 stack");
6101 }
6102
6103 /// A `long double` on its own is a different answer, because on its own it comes back on the
6104 /// x87 stack rather than in a register, which is somewhere the call cannot be said to write.
6105 ///
6106 /// So the call gives back nothing at all and the value is taken off the stack by the `fstp`
6107 /// straight after it. That instruction has to be straight after it: the stack is one place and
6108 /// anything else that touched it before this ran would be looking at the value still on it.
6109 #[test]
6110 fn a_call_that_gives_back_a_long_double_takes_it_off_the_stack_at_once() {
6111 let (mut names, mut source, block, _) = blank(&[]);
6112 let long_double = Type::float(rucc_ir::Float::F80);
6113 let sig = source.add_signature(Signature::new().with_returns(&[long_double]));
6114 let callee = names.intern("g");
6115 Builder::new(&mut source, block).call(callee, sig, &[]);
6116
6117 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
6118 .expect("the value comes back in st0");
6119 let text = mir::print_func(&lowered.func, &names, ®S);
6120 let after: Vec<&str> =
6121 text.lines().skip_while(|line| !line.contains("x64.call")).skip(1).collect();
6122 assert_eq!(after[0].trim(), "%0:gpr = x64.lea_64 [$rsp]", "{text}");
6123 assert_eq!(after[1].trim(), "x64.fstp_t [%0]", "{text}");
6124 // And the slot it went into is the sixteen bytes the type takes, like every other one.
6125 assert_eq!(lowered.stack.locals.len(), 1, "{text}");
6126 assert_eq!(lowered.stack.locals[0].size, X87_BYTES);
6127 }
6128
6129 #[test]
6130 fn a_call_through_an_address_goes_through_the_register_the_address_is_in() {
6131 let i32 = Type::int(32);
6132 let (mut names, mut source, block, args) = blank(&[Type::PTR, i32]);
6133 let sig = source.add_signature(Signature::new().with_params(&[i32]).with_returns(&[i32]));
6134 let varargs = source.push_abis(&[]);
6135 let info = source.add_call(CallInfo { callee: None, signature: sig, varargs });
6136 let mut build = Builder::new(&mut source, block);
6137 let inst = InstData {
6138 args: build.func().push_values(&[args[0], args[1]]),
6139 extra: Extra::Call(info),
6140 ..InstData::new(Opcode::CallIndirect)
6141 };
6142 let called = build.inst(inst, &[i32]);
6143 let got = source[called].first_result.expect("an integer comes back");
6144 Builder::new(&mut source, block).ret(&[got]);
6145
6146 // `int f(int (*g)(int), int a) { return g(a); }`. The first operand is the address and
6147 // the arguments are the ones behind it, and everything else about the call is what a call
6148 // to a name would have been.
6149 let text = lower(&mut names, &source);
6150 assert!(text.contains("= x64.call_reg %0, %1($rdi)"), "{text}");
6151 assert!(text.contains("x64.ret_val_32 %2($rax)"), "{text}");
6152 assert!(!text.contains("@g"), "a call through an address names nobody: {text}");
6153 }
6154
6155 #[test]
6156 fn an_instruction_no_rule_covers_is_reported() {
6157 let (mut names, mut source, block, args) = blank(&[Type::PTR]);
6158 let mut build = Builder::new(&mut source, block);
6159 let operands = build.func().push_values(&[args[0]]);
6160 build.inst(InstData { args: operands, ..InstData::new(Opcode::MetaBegin) }, &[]);
6161
6162 // The mark that an object has come into being, which nothing writes an instruction for
6163 // yet: what it needs is a write over a range of the lifetime plane, and that is
6164 // `tamnd/rucc#856`. Nothing about it is a width or a register, so there is nothing for the
6165 // message to add beyond the name.
6166 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
6167 .expect_err("no rule writes the beginning of a lifetime");
6168 assert_eq!(failed.to_string(), "no rule lowers a `meta_begin`");
6169
6170 // It produces nothing, so there is no type in the message and nothing invents one, and the
6171 // instruction comes back so a caller can ask the function where it was.
6172 let inst = failed.inst().expect("the instruction it is about");
6173 assert_eq!(source[inst].opcode, Opcode::MetaBegin);
6174 }
6175
6176 /// A barrier is written by name here, and what it is depends on the ordering and on nothing
6177 /// else. `crate::expand` is where the reasoning about this machine's memory model lives.
6178 #[test]
6179 fn a_barrier_is_one_instruction_at_the_strongest_ordering_and_none_below_it() {
6180 for order in MemOrder::all().filter(|&order| order != MemOrder::NotAtomic) {
6181 let (mut names, mut source, block, _) = blank(&[]);
6182 let mut build = Builder::new(&mut source, block);
6183 build
6184 .inst(InstData { extra: Extra::Order(order), ..InstData::new(Opcode::Fence) }, &[]);
6185
6186 let text = lower(&mut names, &source);
6187 assert_eq!(text.contains("x64.mfence"), order == MemOrder::SeqCst, "{order:?}: {text}");
6188 }
6189 }
6190
6191 /// A compare and exchange is written by name too, and at the width of the value rather than at
6192 /// the width of the address, which is the mistake worth pinning: everything here is a pointer
6193 /// and only the value says how many bytes the instruction touches.
6194 #[test]
6195 fn a_compare_and_exchange_is_one_instruction_at_the_width_of_the_value() {
6196 for bits in [8, 16, 32, 64] {
6197 let ty = Type::int(bits);
6198 let (mut names, mut source, block, args) = blank(&[Type::PTR, ty, ty]);
6199 let mut build = Builder::new(&mut source, block);
6200 let mem = build.func().add_mem(MemInfo {
6201 size: u64::from(bits / 8),
6202 align: bits / 8,
6203 order: MemOrder::SeqCst,
6204 ..plain()
6205 });
6206 let operands = build.func().push_values(&[args[0], args[1], args[2]]);
6207 build.inst(
6208 InstData {
6209 args: operands,
6210 extra: Extra::Mem(mem),
6211 ..InstData::new(Opcode::Cmpxchg)
6212 },
6213 &[ty, Type::I1],
6214 );
6215
6216 // Two values out of one instruction, the first of them in the register the machine
6217 // reads the expected value out of, the second free for the allocator to place. The
6218 // address is the memory operand and neither of the two values is.
6219 let text = lower(&mut names, &source);
6220 let written = format!("%3:gpr($rax), %4:gpr = x64.cmpxchg_{bits} %1($rax), %2, [%0]");
6221 assert!(text.contains(&written), "{bits}: {text}");
6222 }
6223 }
6224
6225 #[test]
6226 fn more_values_back_than_the_convention_has_registers_for_is_reported() {
6227 let i64 = Type::int(64);
6228 let (mut names, mut source, block, args) = blank(&[i64, i64, i64]);
6229 let mut build = Builder::new(&mut source, block);
6230 build.ret(&[args[0], args[1], args[2]]);
6231
6232 // Two integers come back in `rax` and `rdx` and a third has nowhere to go, which is not a
6233 // gap in the rules but the convention saying no. The front end classifies before it gets
6234 // here, so this is the shape that would mean the classification went wrong.
6235 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
6236 .expect_err("only two come back");
6237 assert_eq!(
6238 failed.to_string(),
6239 "what this function gives back takes more registers than this convention has for it"
6240 );
6241
6242 let inst = failed.inst().expect("the instruction it is about");
6243 assert_eq!(source[inst].opcode, Opcode::Return);
6244 }
6245
6246 /// A refusal about a signature has no instruction, which is what makes it the one arm apart.
6247 ///
6248 /// Everything else is about something written somewhere in the body and hands it back so a
6249 /// caller can ask the function where it came from. A parameter arrives before the first
6250 /// instruction runs, so there is nothing in the body to point at and the message is about
6251 /// the function.
6252 #[test]
6253 fn a_refusal_about_a_parameter_has_no_instruction_to_point_at() {
6254 let missing = Unsupported::Argument { index: 0, missing: Missing::OnX87 };
6255 assert_eq!(missing.inst(), None);
6256 }
6257
6258 /// An `alloca` of a fixed size, which is what every local whose address is taken becomes.
6259 fn slot(source: &mut Func, block: Block, size: u64, align: u32) -> Value {
6260 let info = MemInfo { size, align, ..plain() };
6261 let mut build = Builder::new(source, block);
6262 let mem = build.func().add_mem(info);
6263 build.value(InstData { extra: Extra::Mem(mem), ..InstData::new(Opcode::Alloca) }, Type::PTR)
6264 }
6265
6266 #[test]
6267 fn a_local_is_memory_in_the_frame_and_one_instruction_that_says_where() {
6268 let (mut names, mut source, block, _) = blank(&[]);
6269 let slot = slot(&mut source, block, 4, 4);
6270 let mut build = Builder::new(&mut source, block);
6271 let nine = build.iconst(Type::int(32), 9);
6272 build.store(nine, slot, plain(), Flags::default());
6273 let loaded = build.load(Type::int(32), slot, plain(), Flags::default());
6274 build.ret(&[loaded]);
6275
6276 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
6277 .expect("every instruction has a rule");
6278
6279 // Four bytes on the list the frame is laid out from, and the one instruction that reads
6280 // where they went. Its displacement is nothing here because there is no frame yet, and
6281 // which instruction is waiting for which local is what `finish` is handed.
6282 assert_eq!(lowered.stack.locals, vec![Local { size: 4, align: 4 }]);
6283 assert_eq!(lowered.stack.addresses.len(), 1);
6284 assert_eq!(lowered.stack.addresses[0].1, 0);
6285 assert_eq!(
6286 mir::print_func(&lowered.func, &names, ®S),
6287 "mfunc @f {\nblock0:\n %0:gpr = x64.lea_64 [$rsp]\n \
6288 %1:gpr = x64.mov_ri_32 9\n x64.mov_mr_32 %1, [%0]\n \
6289 %2:gpr = x64.mov_rm_32 [%0]\n x64.ret_val_32 %2($rax)\n}\n"
6290 );
6291 }
6292
6293 #[test]
6294 fn a_local_the_program_declared_says_which_declaration_it_is_and_the_rest_say_nothing() {
6295 let (mut names, mut source, block, _) = blank(&[]);
6296 let scratch = slot(&mut source, block, 4, 4);
6297 let mut build = Builder::new(&mut source, block);
6298 let mem = build.func().add_mem(MemInfo { size: 8, align: 8, ..plain() });
6299 let declared = build
6300 .value(InstData { extra: Extra::Mem(mem), ..InstData::new(Opcode::Alloca) }, Type::PTR);
6301 build.func().declare_mem(mem, 41);
6302 build.store(scratch, declared, MemInfo { size: 8, align: 8, ..plain() }, Flags::default());
6303 build.ret(&[]);
6304
6305 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
6306 .expect("every instruction has a rule");
6307
6308 // Two locals and one declaration, held against the order the allocas were lowered in,
6309 // which is the only name a local has by the time the frame places it. The scratch one was
6310 // reached first and is local zero, so the declared one is local one.
6311 assert_eq!(lowered.stack.locals.len(), 2);
6312 assert_eq!(lowered.stack.declared, vec![(1, 41)]);
6313 }
6314
6315 /// A local the program kept in a value comes out saying which register holds it.
6316 ///
6317 /// The other half of the local above, which had a slot. This one has none, so what carries the
6318 /// declaration is the register the instruction computing it writes into.
6319 #[test]
6320 fn a_local_the_program_kept_in_a_value_says_which_register_holds_it() {
6321 let (mut names, mut source, block, _) = blank(&[]);
6322 let mut build = Builder::new(&mut source, block);
6323 let nine = build.iconst(Type::int(32), 9);
6324 let ten = build.iconst(Type::int(32), 10);
6325 let sum = build.binary(Opcode::Add, nine, ten, Flags::default());
6326 build.func().declare_value(sum, 41);
6327 build.ret(&[sum]);
6328
6329 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
6330 .expect("every instruction has a rule");
6331
6332 // One pair and not three. The constants are values the program never declared, and a
6333 // register holding one of those is nobody's. The register is the one the addition writes,
6334 // which the listing under it is what pins down.
6335 assert_eq!(lowered.func.named, vec![(41, mir::Reg::virtual_reg(1))]);
6336 assert_eq!(
6337 mir::print_func(&lowered.func, &names, ®S),
6338 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_ri_32 9\n \
6339 %1:gpr(reuse 1) = x64.add_ri_32 %0, 10\n x64.ret_val_32 %1($rax)\n}\n"
6340 );
6341 }
6342
6343 /// A local held in a constant two blocks want is two registers and both of them are it.
6344 ///
6345 /// Why the declaration is written down as each register is handed out rather than once at the
6346 /// end over the map from values to registers. That map remembers the last register a value was
6347 /// written into, and a constant is written again in every block that wants one, so a local held
6348 /// in one would come out findable in the last block of the function and nowhere else.
6349 #[test]
6350 fn a_local_held_in_a_constant_two_blocks_want_is_named_in_both_of_them() {
6351 let i32 = Type::int(32);
6352 let (mut names, mut source, entry, args) = blank(&[i32, i32]);
6353 let then = source.create_block();
6354 let other = source.create_block();
6355 let join = source.create_block();
6356 let got = source.append_param(join, i32);
6357
6358 let mut build = Builder::new(&mut source, entry);
6359 let seven = build.iconst(i32, 7);
6360 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
6361 build.func().declare_value(seven, 41);
6362 build.br_if(cond, then, &[], other, &[]);
6363 Builder::new(&mut source, then).jump(join, &[seven]);
6364 Builder::new(&mut source, other).jump(join, &[seven]);
6365 Builder::new(&mut source, join).ret(&[got]);
6366
6367 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
6368 .expect("every instruction has a rule");
6369
6370 let held = &lowered.func.named;
6371 assert_eq!(held.len(), 2, "one register per block that wanted the seven: {held:?}");
6372 assert!(held.iter().all(|&(decl, _)| decl == 41), "{held:?}");
6373 assert_ne!(held[0].1, held[1].1, "the same register in two blocks: {held:?}");
6374 }
6375
6376 /// A parameter the program declared comes out named too, in the register it arrived in.
6377 ///
6378 /// The case the walk over the map at the end is for. A parameter is put in a register the
6379 /// convention chose rather than in a fresh one, so nothing asks the mint for it and the pair
6380 /// would otherwise never be written down.
6381 #[test]
6382 fn a_parameter_the_program_declared_says_which_register_it_arrived_in() {
6383 let i32 = Type::int(32);
6384 let (mut names, mut source, block, args) = blank(&[i32]);
6385 let mut build = Builder::new(&mut source, block);
6386 build.func().declare_value(args[0], 41);
6387 build.ret(&[args[0]]);
6388
6389 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
6390 .expect("every instruction has a rule");
6391
6392 let held = &lowered.func.named;
6393 assert_eq!(held.len(), 1, "one pair for the one parameter: {held:?}");
6394 assert_eq!(held[0].0, 41);
6395 }
6396
6397 /// A function with nothing declared in it says nothing, which is every function compiled
6398 /// without debugging information asked for.
6399 #[test]
6400 fn a_function_the_front_end_named_nothing_in_names_no_registers() {
6401 let (mut names, mut source, block, _) = blank(&[]);
6402 let mut build = Builder::new(&mut source, block);
6403 let nine = build.iconst(Type::int(32), 9);
6404 build.ret(&[nine]);
6405
6406 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
6407 .expect("every instruction has a rule");
6408 assert!(lowered.func.named.is_empty(), "{:?}", lowered.func.named);
6409 }
6410
6411 #[test]
6412 fn the_frame_is_what_fills_the_address_of_a_local_in() {
6413 let (mut names, mut source, block, _) = blank(&[]);
6414 let slot = slot(&mut source, block, 4, 4);
6415 let mut build = Builder::new(&mut source, block);
6416 let nine = build.iconst(Type::int(32), 9);
6417 build.store(nine, slot, plain(), Flags::default());
6418 let loaded = build.load(Type::int(32), slot, plain(), Flags::default());
6419 build.ret(&[loaded]);
6420
6421 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
6422 .expect("every instruction has a rule");
6423 let stack = lowered.stack;
6424 let mut out = lowered.func;
6425 let env = env();
6426 let allocation = rucc_regalloc::run(&mut out, &env, "test", true);
6427 let layout = stack.layout(Layout::new(&SYSV, REGS));
6428 let frame = Frame::of(&out, &allocation, &layout);
6429 finish(&mut out, &allocation, &frame, &stack, Convention::new(&SYSV, &FRAME), &mut names);
6430
6431 // `int f(void) { int x; x = 9; return x; }` with the address of `x` taken, end to end.
6432 // A leaf small enough to live in the red zone takes no frame at all, so the stack pointer
6433 // never moves and the four bytes are below it, which is what the negative offset is. The
6434 // instruction the lowering left with nothing in its displacement now has the answer in it.
6435 let text = mir::print_func(&out, &names, ®S);
6436 assert!(text.contains("$rax = x64.lea_64 [$rsp - 8]"), "{text}");
6437 assert!(!text.contains("x64.sub_ri_64"), "{text}");
6438 assert_eq!(frame.size(), 0);
6439 assert_eq!(frame.local(0), Some(-8));
6440 }
6441
6442 /// An `alloca` whose size is an operand, which is a variable length array.
6443 fn growing(source: &mut Func, block: Block, size: Value, align: u32) -> Value {
6444 let info = MemInfo { size: 0, align, ..plain() };
6445 let mut build = Builder::new(source, block);
6446 let mem = build.func().add_mem(info);
6447 let args = build.func().push_values(&[size]);
6448 build.value(
6449 InstData { args, extra: Extra::Mem(mem), ..InstData::new(Opcode::Alloca) },
6450 Type::PTR,
6451 )
6452 }
6453
6454 #[test]
6455 fn a_stack_slot_whose_size_is_not_known_until_it_runs_takes_the_bytes_off_the_stack_pointer() {
6456 let (mut names, mut source, block, args) = blank(&[Type::int(64)]);
6457 let slot = growing(&mut source, block, args[0], 16);
6458 Builder::new(&mut source, block).ret(&[slot]);
6459
6460 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
6461 .expect("every instruction has a rule");
6462
6463 // The bytes come off the stack pointer where the declaration stands and the address is
6464 // where the stack pointer then is, which is one subtraction and one `lea` rather than a
6465 // slot the frame laid out. Nothing is on the list of locals, because there is nothing
6466 // about this the frame could place.
6467 let text = mir::print_func(&lowered.func, &names, ®S);
6468 assert!(text.contains("$rsp = x64.sub_rr_64 $rsp, %0"), "{text}");
6469 assert!(text.contains("x64.lea_64 [$rsp]"), "{text}");
6470 assert!(lowered.stack.locals.is_empty(), "{text}");
6471 assert_eq!(lowered.stack.dynamic.len(), 1);
6472 assert!(lowered.stack.grown_at.is_some());
6473 }
6474
6475 #[test]
6476 fn a_growing_slot_wanting_more_alignment_than_the_stack_pointer_has_is_reported() {
6477 let (mut names, mut source, block, args) = blank(&[Type::int(64)]);
6478 let slot = growing(&mut source, block, args[0], 32);
6479 Builder::new(&mut source, block).ret(&[slot]);
6480
6481 // Thirty two is more than a call leaves the stack pointer on, so giving it what it asked
6482 // for means masking the stack pointer after moving it, and after that no constant reaches
6483 // the rest of the frame from the frame pointer either. A second pointer held for the
6484 // purpose is what fixes it and there is not one yet.
6485 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
6486 .expect_err("nothing realigns a frame that grows");
6487 assert_eq!(
6488 failed.to_string(),
6489 "this local wants more alignment than the stack pointer is left on, which needs a \
6490 base register nothing here keeps"
6491 );
6492 }
6493
6494 #[test]
6495 fn a_frame_that_grows_reaches_its_own_locals_through_the_frame_pointer() {
6496 let (mut names, mut source, block, args) = blank(&[Type::int(64)]);
6497 let fixed = slot(&mut source, block, 4, 4);
6498 let mut build = Builder::new(&mut source, block);
6499 let nine = build.iconst(Type::int(32), 9);
6500 build.store(nine, fixed, plain(), Flags::default());
6501 let grown = growing(&mut source, block, args[0], 16);
6502 Builder::new(&mut source, block).ret(&[grown]);
6503
6504 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
6505 .expect("every instruction has a rule");
6506 let stack = lowered.stack;
6507 let mut out = lowered.func;
6508 let env = env();
6509 let allocation = rucc_regalloc::run(&mut out, &env, "test", true);
6510 let layout = stack.layout(Layout::new(&SYSV, REGS));
6511 let frame = Frame::of(&out, &allocation, &layout);
6512 finish(&mut out, &allocation, &frame, &stack, Convention::new(&SYSV, &FRAME), &mut names);
6513
6514 // The stack pointer moves in the middle of the function, so the four bytes of the fixed
6515 // local are not a constant away from it any more and the frame pointer is what reaches
6516 // them. The frame keeps one whatever the flags asked for, takes its bytes rather than
6517 // living in the red zone, and the address of the growing slot is off the stack pointer as
6518 // it stands after the subtraction rather than off anything the prologue left.
6519 let text = mir::print_func(&out, &names, ®S);
6520 assert!(frame.grows());
6521 assert!(frame.frame_pointer());
6522 assert!(frame.size() > 0, "{text}");
6523 assert!(text.contains("x64.lea_64 [$rbp"), "{text}");
6524 assert!(text.contains("$rsp = x64.sub_rr_64 $rsp"), "{text}");
6525 assert!(text.contains("x64.lea_64 [$rsp]"), "{text}");
6526 }
6527
6528 #[test]
6529 fn an_address_is_read_written_and_added_to_like_the_integer_it_is() {
6530 let (mut names, mut source, block, args) = blank(&[Type::PTR, Type::int(64)]);
6531 let mut build = Builder::new(&mut source, block);
6532 let stepped = build.func().push_values(&[args[0], args[1]]);
6533 let next =
6534 build.value(InstData { args: stepped, ..InstData::new(Opcode::PtrAdd) }, Type::PTR);
6535 let loaded = build.load(Type::int(32), next, plain(), Flags::default());
6536 build.ret(&[loaded]);
6537
6538 // `int f(int *p, long i) { return *(int *)((char *)p + i); }`. Nothing about this is new
6539 // in the rule set, which is the point: the two addresses arrive in registers because an
6540 // address is an integer as wide as one, and the arithmetic on them is the add it always
6541 // was, so every rule written about an add reaches it.
6542 //
6543 // The add stays its own instruction here rather than folding into the address the load
6544 // reads from. Two registers with no scale on either is the one addressing mode the rules
6545 // have no load through, because the folds that exist are the displacement one and the
6546 // scaled ones, and this is neither. `crate::fold` is what puts the two together, after
6547 // selection, and this is the pair it is handed.
6548 assert_eq!(
6549 lower(&mut names, &source),
6550 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
6551 %1:gpr($rsi) = x64.arg_val_64\n %2:gpr(reuse 1) = x64.add_rr_64 %0, %1\n \
6552 %3:gpr = x64.mov_rm_32 [%2]\n x64.ret_val_32 %3($rax)\n}\n"
6553 );
6554 }
6555
6556 /// The address of a file scope name, which is what every use of a global and every string
6557 /// literal starts from.
6558 fn address_of(source: &mut Func, block: Block, names: &mut Interner, name: &str) -> Value {
6559 let symbol = names.intern(name);
6560 let mut build = Builder::new(source, block);
6561 build.value(
6562 InstData { extra: Extra::Symbol(symbol), ..InstData::new(Opcode::GlobalAddr) },
6563 Type::PTR,
6564 )
6565 }
6566
6567 #[test]
6568 fn the_address_of_a_name_is_one_instruction_carrying_the_name() {
6569 let (mut names, mut source, block, _) = blank(&[]);
6570 let counter = address_of(&mut source, block, &mut names, "counter");
6571 let mut build = Builder::new(&mut source, block);
6572 let loaded = build.load(Type::int(32), counter, plain(), Flags::default());
6573 build.ret(&[loaded]);
6574
6575 // `extern int counter; int f(void) { return counter; }`. The address is an addressing mode
6576 // that names no register and carries the symbol, which is what the assembler writes
6577 // relative to `%rip` and what the object writer leaves a relocation for.
6578 assert_eq!(
6579 lower(&mut names, &source),
6580 "mfunc @f {\nblock0:\n %0:gpr = x64.lea_64 [@counter]\n \
6581 %1:gpr = x64.mov_rm_32 [%0]\n x64.ret_val_32 %1($rax)\n}\n"
6582 );
6583 }
6584
6585 #[test]
6586 fn the_address_of_a_name_outside_the_file_is_read_out_of_the_offset_table() {
6587 let (mut names, mut source, block, _) = blank(&[]);
6588 let away = address_of(&mut source, block, &mut names, "away");
6589 Builder::new(&mut source, block).ret(&[away]);
6590 let elsewhere: Elsewhere = [names.intern("away")].into_iter().collect();
6591
6592 // `extern void away(void); void *f(void) { return away; }`. A load and not an address
6593 // computation, because the distance from here to a name a shared library may be the one
6594 // that defines is not a number any link can work out, and the slot the linker fills in is
6595 // in this program and so is a distance it has.
6596 let out =
6597 func(&source, &mut names, &SYSV, &elsewhere).expect("every instruction has a rule");
6598 assert_eq!(
6599 mir::print_func(&out.func, &names, ®S),
6600 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_rm_64 [got @away]\n \
6601 x64.ret_val_64 %0($rax)\n}\n"
6602 );
6603 }
6604
6605 #[test]
6606 fn the_address_of_a_thread_local_is_an_offset_out_of_the_table_plus_where_this_thread_starts() {
6607 let (mut names, mut source, block, _) = blank(&[]);
6608 let own = address_of(&mut source, block, &mut names, "own");
6609 Builder::new(&mut source, block).ret(&[own]);
6610 let elsewhere = Elsewhere::default().with_threads([names.intern("own")]);
6611
6612 // `extern _Thread_local int own; void *f(void) { return &own; }`. Three instructions where
6613 // the two cases above are one, because there is no address to load or to work out: the
6614 // slot holds how far into a thread's block the variable sits, `%fs:0` is where this
6615 // thread's block starts, and the sum of the two is this thread's copy.
6616 let out =
6617 func(&source, &mut names, &SYSV, &elsewhere).expect("every instruction has a rule");
6618 assert_eq!(
6619 mir::print_func(&out.func, &names, ®S),
6620 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_rm_64 [thread @own]\n \
6621 %1:gpr = x64.mov_rm_64 [fs:0]\n %2:gpr(reuse 1) = x64.add_rr_64 %0, %1\n \
6622 x64.ret_val_64 %2($rax)\n}\n"
6623 );
6624 }
6625
6626 /// The same load with nothing added to it, which is the whole of `__builtin_thread_pointer`.
6627 #[test]
6628 fn the_start_of_this_thread_s_own_storage_is_the_one_load_and_no_arithmetic() {
6629 let (mut names, mut source, block, _) = blank(&[]);
6630 let here =
6631 Builder::new(&mut source, block).value(InstData::new(Opcode::ThreadPointer), Type::PTR);
6632 Builder::new(&mut source, block).ret(&[here]);
6633
6634 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
6635 .expect("every instruction has a rule");
6636 assert_eq!(
6637 mir::print_func(&out.func, &names, ®S),
6638 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_rm_64 [fs:0]\n \
6639 x64.ret_val_64 %0($rax)\n}\n"
6640 );
6641 }
6642
6643 /// One `asm` statement, with its template and its constraint list written as a program does.
6644 fn assembly(
6645 source: &mut Func,
6646 block: Block,
6647 names: &mut Interner,
6648 template: &str,
6649 constraints: &str,
6650 args: &[Value],
6651 results: &[Type],
6652 ) -> Inst {
6653 clobbering(source, block, names, template, constraints, "memory", args, results)
6654 }
6655
6656 /// The same with a clobber list of its own, for the statements that are about one.
6657 #[allow(clippy::too_many_arguments)]
6658 fn clobbering(
6659 source: &mut Func,
6660 block: Block,
6661 names: &mut Interner,
6662 template: &str,
6663 constraints: &str,
6664 clobbers: &str,
6665 args: &[Value],
6666 results: &[Type],
6667 ) -> Inst {
6668 let info = AsmInfo {
6669 template: names.intern(template),
6670 constraints: names.intern(constraints),
6671 clobbers: names.intern(clobbers),
6672 targets: rucc_ir::BlockCallList::EMPTY,
6673 };
6674 Builder::new(source, block).inline_asm(info, args, results, Flags::VOLATILE)
6675 }
6676
6677 /// What a program asking the processor what it can do writes, which is the instruction whose
6678 /// every operand is a register its text does not name.
6679 #[test]
6680 fn a_template_whose_registers_are_named_by_the_constraints_places_them_from_the_letters() {
6681 let u32 = Type::int(32);
6682 let (mut names, mut source, block, _) = blank(&[]);
6683 let zero = Builder::new(&mut source, block).iconst(u32, 0);
6684 let out = clobbering(
6685 &mut source,
6686 block,
6687 &mut names,
6688 "cpuid",
6689 "=a,a",
6690 "ebx,ecx,edx",
6691 &[zero],
6692 &[u32],
6693 );
6694 let produced = source[out].results().next().expect("one result");
6695 Builder::new(&mut source, block).ret(&[produced]);
6696
6697 // `asm ("cpuid" : "=a" (n) : "a" (0) : "ebx", "ecx", "edx")`, which is the first thing
6698 // every program that has a faster path on some machines writes. Four registers written and
6699 // two read, none of them in the template, all of them out of the description, and the two
6700 // that the letters named are the statement's own. The subleaf is a zero because the
6701 // instruction reads `ecx` and the program said nothing about what is in it. The three
6702 // clobbers are gone because `cpuid` writes those three anyway, and saying it twice is one
6703 // register with two definitions.
6704 assert_eq!(
6705 lower(&mut names, &source),
6706 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_ri_32 0\n \
6707 %1:gpr = x64.mov_ri_64 0\n \
6708 %2:gpr($rax), %3:gpr($rbx), %4:gpr($rcx), %5:gpr($rdx) = x64.cpuid %0($rax), \
6709 %1($rcx)\n x64.ret_val_32 %2($rax)\n}\n"
6710 );
6711 }
6712
6713 /// An operand the program pinned, by declaring the object it comes from `register long x asm
6714 /// ("r12")`. The letter on its own leaves the allocator to pick, and a template that reads the
6715 /// register by name needs the two to be the same register, so the brace is what ties them
6716 /// together. That is the one use of a local register variable the GNU manual calls reliable,
6717 /// and it is what tcc's `tests/tcctest.c` counts on.
6718 #[test]
6719 fn an_operand_the_program_pinned_is_placed_in_the_register_it_named() {
6720 let u64 = Type::int(64);
6721 let (mut names, mut source, block, _) = blank(&[]);
6722 let out =
6723 assembly(&mut source, block, &mut names, "mov $0x4542, %r12", "=r{r12}", &[], &[u64]);
6724 let produced = source[out].results().next().expect("one result");
6725 Builder::new(&mut source, block).ret(&[produced]);
6726
6727 // The template is one instruction the table already has, so it lowers to that instruction
6728 // rather than to text nobody read, and the register it names is the statement's own output
6729 // because the brace put the output there. Without the brace the letter would have let the
6730 // allocator pick, the two `%r12` would have been different registers, and the program would
6731 // have come back with whatever was in the one it picked.
6732 assert_eq!(
6733 lower(&mut names, &source),
6734 "mfunc @f {\nblock0:\n %0:gpr($r12) = x64.mov_ri_64 17730\n \
6735 x64.ret_val_64 %0($rax)\n}\n"
6736 );
6737 }
6738
6739 /// A clobber the instruction does not write itself, which is the case the list is there for.
6740 /// It goes on as a definition of the register, in among the other definitions, because that is
6741 /// the whole of how a machine function says a register is not worth anything after this.
6742 #[test]
6743 fn a_clobber_the_instruction_does_not_write_itself_is_a_definition_of_that_register() {
6744 let (mut names, mut source, block, _) = blank(&[]);
6745 clobbering(&mut source, block, &mut names, "pause", "", "rsi,cc,memory", &[], &[]);
6746 Builder::new(&mut source, block).ret(&[]);
6747
6748 assert_eq!(lower(&mut names, &source), "mfunc @f {\nblock0:\n $rsi = x64.pause\n}\n");
6749 }
6750
6751 /// A clobber naming something this has no register for. Refused rather than dropped, since the
6752 /// list is the program saying which registers it may not leave anything in, and an entry
6753 /// nobody read is a register something may still be left in.
6754 #[test]
6755 fn a_clobber_this_has_no_register_for_is_refused() {
6756 let (mut names, mut source, block, _) = blank(&[]);
6757 clobbering(&mut source, block, &mut names, "pause", "", "zmm0", &[], &[]);
6758 Builder::new(&mut source, block).ret(&[]);
6759
6760 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
6761 .expect_err("there is no such register here");
6762 assert_eq!(
6763 failed.to_string(),
6764 "this `asm` says it destroys a register this has no name for"
6765 );
6766 }
6767
6768 #[test]
6769 fn an_asm_with_an_empty_template_and_no_operands_is_no_instructions() {
6770 let (mut names, mut source, block, _) = blank(&[]);
6771 assembly(&mut source, block, &mut names, "", "", &[], &[]);
6772 Builder::new(&mut source, block).ret(&[]);
6773
6774 // `asm volatile ("" : : : "memory")`, which is a barrier and nothing else. The barrier was
6775 // spent on the optimizer, which has finished by now, so what is left is nothing.
6776 assert_eq!(lower(&mut names, &source), "mfunc @f {\nblock0:\n}\n");
6777 }
6778
6779 #[test]
6780 fn an_output_an_input_is_tied_to_is_the_register_that_input_arrived_in() {
6781 let i32 = Type::int(32);
6782 let (mut names, mut source, block, args) = blank(&[i32]);
6783 let out = assembly(&mut source, block, &mut names, "", "=r,0", &args, &[i32]);
6784 let produced = source[out].results().next().expect("one result");
6785 Builder::new(&mut source, block).ret(&[produced]);
6786
6787 // `asm ("" : "=r" (x) : "0" (x))`, which is how a program stops the optimizer following a
6788 // value without changing it. The two share a place and the template writes nothing over
6789 // it, so the value comes back out of the register it went in.
6790 assert_eq!(
6791 lower(&mut names, &source),
6792 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
6793 x64.ret_val_32 %0($rax)\n}\n"
6794 );
6795 }
6796
6797 #[test]
6798 fn an_output_written_plus_is_the_same_rename() {
6799 let i32 = Type::int(32);
6800 let (mut names, mut source, block, args) = blank(&[i32]);
6801 let out = assembly(&mut source, block, &mut names, "", "+r", &args, &[i32]);
6802 let produced = source[out].results().next().expect("one result");
6803 Builder::new(&mut source, block).ret(&[produced]);
6804
6805 // `asm ("" : "+r" (x))`, which says the same thing in one operand instead of two.
6806 assert_eq!(
6807 lower(&mut names, &source),
6808 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
6809 x64.ret_val_32 %0($rax)\n}\n"
6810 );
6811 }
6812
6813 #[test]
6814 fn an_output_nothing_is_tied_to_is_a_zero() {
6815 let i32 = Type::int(32);
6816 let (mut names, mut source, block, _) = blank(&[]);
6817 let out = assembly(&mut source, block, &mut names, "", "=r", &[], &[i32]);
6818 let produced = source[out].results().next().expect("one result");
6819 Builder::new(&mut source, block).ret(&[produced]);
6820
6821 // `asm ("" : "=r" (y))`, whose answer is whatever the assembly left in the register, and
6822 // an empty template leaves nothing. A definite value rather than a register nothing wrote,
6823 // because the allocator is owed a definition before the use however little the program is.
6824 assert_eq!(
6825 lower(&mut names, &source),
6826 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_ri_32 0\n x64.ret_val_32 %0($rax)\n}\n"
6827 );
6828 }
6829
6830 #[test]
6831 fn a_template_that_is_one_instruction_becomes_that_instruction() {
6832 let (mut names, mut source, block, _) = blank(&[]);
6833 assembly(&mut source, block, &mut names, "pause", "", &[], &[]);
6834 Builder::new(&mut source, block).ret(&[]);
6835
6836 // `asm volatile ("pause")`, which is what every spin lock in every allocator writes. One
6837 // instruction, no operands, and nothing between the template and the machine but the table
6838 // that already says what a `pause` is.
6839 assert_eq!(lower(&mut names, &source), "mfunc @f {\nblock0:\n x64.pause\n}\n");
6840 }
6841
6842 #[test]
6843 fn a_template_that_reads_a_segment_becomes_the_load_it_already_was() {
6844 let i64 = Type::int(64);
6845 let (mut names, mut source, block, _) = blank(&[]);
6846 let out = assembly(&mut source, block, &mut names, "movq %%fs:0, %0", "=r", &[], &[i64]);
6847 let produced = source[out].results().next().expect("one result");
6848 Builder::new(&mut source, block).ret(&[produced]);
6849
6850 // `asm ("movq %%fs:0, %0" : "=r" (tid))`, which is how a program finds the block its own
6851 // thread owns. The same instruction `crate::lower` already writes for a thread-local
6852 // variable, reached this time because a program wrote it out by hand.
6853 assert_eq!(
6854 lower(&mut names, &source),
6855 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_rm_64 [fs:0]\n \
6856 x64.ret_val_64 %0($rax)\n}\n"
6857 );
6858 }
6859
6860 /// A template this cannot read is kept as its text, which is what gcc does with every template.
6861 /// Whether the text is an instruction is the assembler's question, asked when the unit is
6862 /// assembled from its listing.
6863 #[test]
6864 fn a_template_naming_an_instruction_this_machine_has_not_got_is_kept_as_text() {
6865 let (mut names, mut source, block, _) = blank(&[]);
6866 assembly(&mut source, block, &mut names, "hcf", "", &[], &[]);
6867 Builder::new(&mut source, block).ret(&[]);
6868
6869 let printed = lower(&mut names, &source);
6870 assert!(printed.contains("x64.template"), "{printed}");
6871 assert!(printed.contains("@hcf"), "{printed}");
6872 }
6873
6874 /// A template kept as text with an operand in a register is refused, since nothing here spells
6875 /// a register into the text yet, and the refusal is about the template.
6876 #[test]
6877 fn a_template_kept_as_text_with_an_operand_in_a_register_is_refused() {
6878 let i32 = Type::int(32);
6879 let (mut names, mut source, block, args) = blank(&[i32]);
6880 assembly(&mut source, block, &mut names, "hcf %0", "r", &[args[0]], &[]);
6881 Builder::new(&mut source, block).ret(&[]);
6882
6883 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
6884 .expect_err("a register is not spelled into kept text");
6885 assert_eq!(
6886 failed.to_string(),
6887 "this `asm` has instructions in its template, which nothing here assembles"
6888 );
6889 }
6890
6891 /// A register the template named is placed as itself, fixed to the register the program wrote
6892 /// down. A register a constraint letter names is a different thing and is placed too, which the
6893 /// test above is about: there the statement said which of its own operands is in the register,
6894 /// and a name in the middle of a template says the register and nothing about any operand.
6895 #[test]
6896 fn a_template_naming_a_register_gets_that_register() {
6897 let i64 = Type::int(64);
6898 let (mut names, mut source, block, _) = blank(&[]);
6899 let out = assembly(&mut source, block, &mut names, "movq %%rax, %0", "=r", &[], &[i64]);
6900 let produced = source[out].results().next().expect("one result");
6901 Builder::new(&mut source, block).ret(&[produced]);
6902
6903 // `asm ("movq %%rax, %0" : "=r" (x))`, which is a program reading whatever is in `%rax`.
6904 // The source is the register itself and the destination is one the allocator picks.
6905 assert_eq!(
6906 lower(&mut names, &source),
6907 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_rr_64 $rax($rax)\n \
6908 x64.ret_val_64 %0($rax)\n}\n"
6909 );
6910 }
6911
6912 /// The half of the same thing every register saving template needs. micropython writes the
6913 /// callee-saved registers into a buffer one `movq %%r12, 48(%%rdi)` at a time, and both halves
6914 /// of that line are a register the template named: the one being stored and the one the address
6915 /// is counted from.
6916 #[test]
6917 fn a_template_counting_an_address_from_a_register_it_named_gets_that_register() {
6918 let (mut names, mut source, block, _) = blank(&[]);
6919 assembly(&mut source, block, &mut names, "movq %%r12, 48(%%rdi)", "", &[], &[]);
6920 Builder::new(&mut source, block).ret(&[]);
6921
6922 assert_eq!(
6923 lower(&mut names, &source),
6924 "mfunc @f {\nblock0:\n x64.mov_mr_64 $r12($r12), [$rdi + 48]\n}\n"
6925 );
6926 }
6927
6928 /// A local kept in a named register, which is the same register named as itself and reached
6929 /// from the other side. micropython's collector writes six of these and reads them with
6930 /// ordinary C rather than with a template.
6931 #[test]
6932 fn a_local_kept_in_a_named_register_is_one_move_out_of_it() {
6933 let (mut names, mut source, block, _) = blank(&[]);
6934 let held = names.intern("rbx");
6935 let value = Builder::new(&mut source, block).value(
6936 InstData { extra: Extra::Symbol(held), ..InstData::new(Opcode::RegisterValue) },
6937 Type::int(64),
6938 );
6939 Builder::new(&mut source, block).ret(&[value]);
6940
6941 assert_eq!(
6942 lower(&mut names, &source),
6943 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_rr_64 $rbx($rbx)\n \
6944 x64.ret_val_64 %0($rax)\n}\n"
6945 );
6946 }
6947
6948 /// The sigil gcc allows in front of the name is syntax and comes off, and a name that is not
6949 /// a register of this machine is refused in words that say which name it was.
6950 #[test]
6951 fn a_register_name_is_read_with_or_without_its_sigil_and_refused_when_there_is_no_such_one() {
6952 for written in ["%r12", "r12"] {
6953 let (mut names, mut source, block, _) = blank(&[]);
6954 let held = names.intern(written);
6955 let value = Builder::new(&mut source, block).value(
6956 InstData { extra: Extra::Symbol(held), ..InstData::new(Opcode::RegisterValue) },
6957 Type::int(64),
6958 );
6959 Builder::new(&mut source, block).ret(&[value]);
6960 assert!(lower(&mut names, &source).contains("$r12($r12)"), "{written} is not read");
6961 }
6962
6963 let (mut names, mut source, block, _) = blank(&[]);
6964 let held = names.intern("nowhere");
6965 let value = Builder::new(&mut source, block).value(
6966 InstData { extra: Extra::Symbol(held), ..InstData::new(Opcode::RegisterValue) },
6967 Type::int(64),
6968 );
6969 Builder::new(&mut source, block).ret(&[value]);
6970
6971 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
6972 .expect_err("there is no such register");
6973 assert_eq!(
6974 failed.to_string(),
6975 "this object is kept in `nowhere`, which is not a register this machine has"
6976 );
6977 }
6978
6979 #[test]
6980 fn a_constraint_list_that_does_not_describe_the_operands_is_refused() {
6981 let i32 = Type::int(32);
6982 let (mut names, mut source, block, args) = blank(&[i32]);
6983 assembly(&mut source, block, &mut names, "", "=r", &args, &[]);
6984 Builder::new(&mut source, block).ret(&[]);
6985
6986 // An output with no result to be, which is what the front end never writes and what a
6987 // hand written module can. Refused rather than placed by a guess.
6988 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
6989 .expect_err("the list and the instruction disagree");
6990 assert_eq!(failed.to_string(), "this `asm` has an operand this cannot place");
6991 }
6992
6993 /// A cast between a pointer and an integer, at whatever width the result is asked for.
6994 fn cast(source: &mut Func, block: Block, opcode: Opcode, from: Value, to: Type) -> Value {
6995 let mut build = Builder::new(source, block);
6996 let args = build.func().push_values(&[from]);
6997 build.value(InstData { args, ..InstData::new(opcode) }, to)
6998 }
6999
7000 #[test]
7001 fn a_cast_between_a_pointer_and_an_integer_as_wide_is_no_instruction_at_all() {
7002 let (mut names, mut source, block, args) = blank(&[Type::PTR]);
7003 let number = cast(&mut source, block, Opcode::PtrToInt, args[0], Type::int(64));
7004 Builder::new(&mut source, block).ret(&[number]);
7005
7006 // `long f(void *p) { return (long)p; }`. An address on this machine is an integer as wide
7007 // as the machine addresses, so the cast changes what the type system calls the value and
7008 // changes nothing about the value, and the register holding it is the one that held it.
7009 assert_eq!(
7010 lower(&mut names, &source),
7011 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
7012 x64.ret_val_64 %0($rax)\n}\n"
7013 );
7014 }
7015
7016 #[test]
7017 fn a_null_pointer_is_a_constant_that_reaches_a_register_before_anything_reads_it() {
7018 let (mut names, mut source, block, _) = blank(&[]);
7019 let mut build = Builder::new(&mut source, block);
7020 let zero = build.iconst(Type::int(64), 0);
7021 let null = cast(&mut source, block, Opcode::IntToPtr, zero, Type::PTR);
7022 Builder::new(&mut source, block).ret(&[null]);
7023
7024 // `void *f(void) { return 0; }`. The cast is nothing, and reading its operand is what
7025 // writes the zero down: a constant is materialized where it is wanted rather than where
7026 // the IR defined it, and without the read there would be no instruction at all.
7027 assert_eq!(
7028 lower(&mut names, &source),
7029 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_ri_64 0\n x64.ret_val_64 %0($rax)\n}\n"
7030 );
7031 }
7032
7033 #[test]
7034 fn the_five_linkages_the_ir_has_narrow_to_the_three_an_object_file_can_say() {
7035 let readings = [
7036 (Linkage::External, mir::Binding::Global),
7037 (Linkage::Common, mir::Binding::Global),
7038 (Linkage::Internal, mir::Binding::Local),
7039 (Linkage::Weak, mir::Binding::Weak),
7040 (Linkage::LinkOnce, mir::Binding::Weak),
7041 ];
7042 for (linkage, wanted) in readings {
7043 let (mut names, mut source, block, _) = blank(&[]);
7044 source.linkage = linkage;
7045 Builder::new(&mut source, block).ret(&[]);
7046 let out = func(&source, &mut names, &SYSV, &Elsewhere::default()).expect("a return");
7047 // The narrowing is done here rather than where the object is written, because a
7048 // machine function is all the assembler and the writer are ever handed.
7049 assert_eq!(out.func.binding, wanted, "{linkage:?}");
7050 }
7051 }
7052
7053 /// The visibility makes the same trip and is not narrowed on the way, because ELF says all
7054 /// three of them.
7055 ///
7056 /// Here for the reason the linkage above is here. A machine function is the whole of what the
7057 /// assembler and the object writer are handed, so a fact about the symbol that does not get
7058 /// onto one is a fact that is gone by the time anything could write it down, and the way that
7059 /// shows up is a shared library exporting the wrong set of names with nothing said anywhere.
7060 #[test]
7061 fn the_visibility_survives_the_trip_from_the_ir_to_a_machine_function() {
7062 let readings = [
7063 (Visibility::Default, mir::Visibility::Default),
7064 (Visibility::Hidden, mir::Visibility::Hidden),
7065 (Visibility::Protected, mir::Visibility::Protected),
7066 ];
7067 for (visibility, wanted) in readings {
7068 let (mut names, mut source, block, _) = blank(&[]);
7069 source.visibility = visibility;
7070 Builder::new(&mut source, block).ret(&[]);
7071 let out = func(&source, &mut names, &SYSV, &Elsewhere::default()).expect("a return");
7072 assert_eq!(out.func.visibility, wanted, "{visibility:?}");
7073 }
7074 }
7075
7076 #[test]
7077 fn a_cast_between_a_pointer_and_a_narrower_integer_is_reported() {
7078 let (mut names, mut source, block, args) = blank(&[Type::PTR]);
7079 let number = cast(&mut source, block, Opcode::PtrToInt, args[0], Type::int(32));
7080 Builder::new(&mut source, block).ret(&[number]);
7081
7082 // The front end never writes one: it casts at the address width and truncates or extends
7083 // around it, so both of those are the rules they always were. IR from somewhere else that
7084 // does write one is refused rather than compiled to a move that keeps the high half.
7085 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
7086 .expect_err("no rule narrows an address");
7087 assert_eq!(failed.to_string(), "no rule lowers a `ptrtoint` producing a `i32`");
7088 }
7089
7090 /// The type this machine has no register for.
7091 fn long_double() -> Type {
7092 Type::float(rucc_ir::Float::F80)
7093 }
7094
7095 #[test]
7096 fn a_double_widened_and_narrowed_again_goes_out_through_the_frame_and_back() {
7097 let f64 = Type::float(rucc_ir::Float::F64);
7098 let (mut names, mut source, block, args) = blank(&[f64]);
7099 let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
7100 let back = cast(&mut source, block, Opcode::FPTrunc, wide, f64);
7101 Builder::new(&mut source, block).ret(&[back]);
7102
7103 // `double f(double d) { long double x = d; return x; }`. The x87 reads memory and nothing
7104 // else, so the value is written to the crossing slot, loaded at the format that widens it
7105 // and put in the slot the eighty bit value lives in. Coming back is the same three the
7106 // other way. Both slots are addressed by a `lea` with nothing in it yet, which is what
7107 // every address in a frame looks like here until `finish` has the numbers.
7108 assert_eq!(
7109 lower(&mut names, &source),
7110 "mfunc @f {\nblock0:\n \
7111 %0:xmm($xmm0) = x64.arg_val_f64\n \
7112 %1:gpr = x64.lea_64 [$rsp]\n \
7113 %2:gpr = x64.lea_64 [$rsp]\n \
7114 x64.movsd_mr %0, [%1]\n \
7115 x64.fld_l [%1]\n \
7116 x64.fstp_t [%2]\n \
7117 %3:gpr = x64.lea_64 [$rsp]\n \
7118 %4:gpr = x64.lea_64 [$rsp]\n \
7119 x64.fld_t [%3]\n \
7120 x64.fstp_l [%4]\n \
7121 %5:xmm = x64.movsd_rm [%4]\n \
7122 x64.ret_val_f64 %5($xmm0)\n}\n"
7123 );
7124 }
7125
7126 #[test]
7127 fn a_long_double_has_sixteen_bytes_of_its_own_and_keeps_them() {
7128 let f64 = Type::float(rucc_ir::Float::F64);
7129 let (mut names, mut source, block, args) = blank(&[f64]);
7130 let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
7131 let once = cast(&mut source, block, Opcode::FPTrunc, wide, f64);
7132 let twice = cast(&mut source, block, Opcode::FPTrunc, wide, f64);
7133 let mut build = Builder::new(&mut source, block);
7134 let sum = build.binary(Opcode::FAdd, once, twice, Flags::default());
7135 build.ret(&[sum]);
7136
7137 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
7138 .expect("every instruction is written");
7139
7140 // Two slots and not four: sixteen bytes for the one eighty bit value, which is what the
7141 // psABI says one takes and is aligned to, and eight for the crossing, which every group
7142 // in the function shares because nothing is ever left in it. The value's slot is its own
7143 // for the whole function, so reading it twice reads the same sixteen bytes.
7144 assert_eq!(
7145 out.stack.locals,
7146 vec![Local { size: 8, align: 8 }, Local { size: 16, align: 16 }]
7147 );
7148 }
7149
7150 #[test]
7151 fn an_integer_becomes_a_long_double_by_being_loaded_as_one() {
7152 let (mut names, mut source, block, args) = blank(&[Type::int(64)]);
7153 let wide = cast(&mut source, block, Opcode::SIToFP, args[0], long_double());
7154 let back =
7155 cast(&mut source, block, Opcode::FPTrunc, wide, Type::float(rucc_ir::Float::F64));
7156 Builder::new(&mut source, block).ret(&[back]);
7157
7158 // `double f(long n) { long double x = n; return x; }`. `fild` is the same push at another
7159 // format, so the conversion is the load and there is no instruction that converts.
7160 let text = lower(&mut names, &source);
7161 assert!(text.contains("x64.mov_mr_64 %0, [%1]"), "{text}");
7162 assert!(text.contains("x64.fild_ll [%1]"), "{text}");
7163 }
7164
7165 #[test]
7166 fn a_long_double_becoming_an_integer_cuts_towards_zero_with_the_control_word() {
7167 let (mut names, mut source, block, args) = blank(&[Type::float(rucc_ir::Float::F64)]);
7168 let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
7169 let whole = cast(&mut source, block, Opcode::FPToSI, wide, Type::int(32));
7170 Builder::new(&mut source, block).ret(&[whole]);
7171
7172 // The one conversion here with no single instruction behind it. C cuts towards zero and
7173 // the unit rounds the way its control word says, so the word is saved, ORed with the two
7174 // bits that mean truncate, loaded, used and put back. Nine instructions for what `fisttp`
7175 // does in one, and `spec/10-backend.md` section 10.8 says why that one is not used.
7176 let text = lower(&mut names, &source);
7177 let group: Vec<&str> = text
7178 .lines()
7179 .map(str::trim)
7180 .filter(|line| line.starts_with("x64.f") || line.contains("_16"))
7181 .collect();
7182 assert_eq!(
7183 group,
7184 [
7185 "x64.fld_l [%1]",
7186 "x64.fstp_t [%2]",
7187 "x64.fnstcw [%5]",
7188 "%6:gpr = x64.mov_rm_16 [%5]",
7189 "%7:gpr(reuse 1) = x64.or_ri_16 %6, 3072",
7190 "x64.mov_mr_16 %7, [%5 + 2]",
7191 "x64.fldcw [%5 + 2]",
7192 "x64.fld_t [%3]",
7193 "x64.fistp_l [%4]",
7194 "x64.fldcw [%5]",
7195 ],
7196 "{text}"
7197 );
7198 }
7199
7200 #[test]
7201 fn a_long_double_is_read_and_written_as_the_bits_it_already_is() {
7202 let (mut names, mut source, block, args) = blank(&[Type::PTR, Type::PTR]);
7203 let mut build = Builder::new(&mut source, block);
7204 let value = build.load(long_double(), args[0], plain(), Flags::default());
7205 build.store(value, args[1], plain(), Flags::default());
7206 build.ret(&[]);
7207
7208 // `void f(long double *a, long double *b) { *b = *a; }`. A copy is a push and a pop at the
7209 // format the value is already in, which neither converts nor looks: a signalling NaN stays
7210 // one and nothing is raised, which is the whole of what makes it a copy.
7211 let text = lower(&mut names, &source);
7212 let group: Vec<&str> =
7213 text.lines().map(str::trim).filter(|line| line.starts_with("x64.f")).collect();
7214 assert_eq!(
7215 group,
7216 ["x64.fld_t [%0]", "x64.fstp_t [%2]", "x64.fld_t [%3]", "x64.fstp_t [%1]"],
7217 "{text}"
7218 );
7219 }
7220
7221 /// Two `long double` values, from two `double` parameters, and the instructions that made
7222 /// them, which every test below this one throws away.
7223 fn two_long_doubles(source: &mut Func, block: Block, args: &[Value]) -> (Value, Value) {
7224 let left = cast(source, block, Opcode::FPExt, args[0], long_double());
7225 let right = cast(source, block, Opcode::FPExt, args[1], long_double());
7226 (left, right)
7227 }
7228
7229 /// The x87 instructions of a function, in order, with everything else dropped.
7230 fn stack_only(text: &str) -> Vec<&str> {
7231 text.lines().map(str::trim).filter(|line| line.contains("x64.f")).collect()
7232 }
7233
7234 /// The two frame slots the last two addresses of a function were taken of, which in a
7235 /// comparison are the two operands in the order they go on the stack.
7236 fn pushed(out: &Lowered) -> Vec<usize> {
7237 let taken: Vec<usize> = out.stack.addresses.iter().map(|&(_, local)| local).collect();
7238 taken[taken.len() - 2..].to_vec()
7239 }
7240
7241 #[test]
7242 fn adding_two_long_doubles_pushes_both_and_leaves_the_answer_in_a_slot() {
7243 let f64 = Type::float(rucc_ir::Float::F64);
7244 let (mut names, mut source, block, args) = blank(&[f64, f64]);
7245 let (left, right) = two_long_doubles(&mut source, block, &args);
7246 let sum =
7247 Builder::new(&mut source, block).binary(Opcode::FAdd, left, right, Flags::default());
7248 let back = cast(&mut source, block, Opcode::FPTrunc, sum, f64);
7249 Builder::new(&mut source, block).ret(&[back]);
7250
7251 // `double f(double a, double b) { return (long double) a + (long double) b; }`. The last
7252 // four lines are the add: both operands pushed, the instruction that names neither of
7253 // them because they are the top two of a stack, and the answer taken off into its slot.
7254 let text = lower(&mut names, &source);
7255 assert_eq!(
7256 stack_only(&text),
7257 [
7258 "x64.fld_l [%2]",
7259 "x64.fstp_t [%3]",
7260 "x64.fld_l [%4]",
7261 "x64.fstp_t [%5]",
7262 "x64.fld_t [%6]",
7263 "x64.fld_t [%7]",
7264 "x64.fadd_p",
7265 "x64.fstp_t [%8]",
7266 "x64.fld_t [%9]",
7267 "x64.fstp_l [%10]",
7268 ],
7269 "{text}"
7270 );
7271 }
7272
7273 #[test]
7274 fn a_subtraction_pushes_the_left_operand_first_and_asks_for_the_att_spelling() {
7275 let f64 = Type::float(rucc_ir::Float::F64);
7276 let (mut names, mut source, block, args) = blank(&[f64, f64]);
7277 let (left, right) = two_long_doubles(&mut source, block, &args);
7278 let less =
7279 Builder::new(&mut source, block).binary(Opcode::FSub, left, right, Flags::default());
7280 let back = cast(&mut source, block, Opcode::FPTrunc, less, f64);
7281 Builder::new(&mut source, block).ret(&[back]);
7282
7283 // The left one goes on first, so it ends up under the right one, and the answer wanted is
7284 // the one below minus the top. In AT&T that is `fsubrp`, since `fsubp` there is `DE E0+i`
7285 // and computes the other one. The `r` says which spelling this is and not which order the
7286 // pushes were in. `crates/rucc/tests/x87.rs` is what says the answer is right, because a
7287 // name is what got this wrong the first time.
7288 let text = lower(&mut names, &source);
7289 assert_eq!(
7290 &stack_only(&text)[4..8],
7291 ["x64.fld_t [%6]", "x64.fld_t [%7]", "x64.fsubr_p", "x64.fstp_t [%8]"],
7292 "{text}"
7293 );
7294 }
7295
7296 #[test]
7297 fn negating_a_long_double_turns_the_sign_over_and_reads_nothing() {
7298 let f64 = Type::float(rucc_ir::Float::F64);
7299 let (mut names, mut source, block, args) = blank(&[f64]);
7300 let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
7301 let flipped = Builder::new(&mut source, block).unary(Opcode::FNeg, wide, long_double());
7302 let back = cast(&mut source, block, Opcode::FPTrunc, flipped, f64);
7303 Builder::new(&mut source, block).ret(&[back]);
7304
7305 // `fchs` and not a subtraction from zero, which would give a different answer at a negative
7306 // zero and would signal at a NaN. It does not read the value as a number at all.
7307 let text = lower(&mut names, &source);
7308 assert_eq!(
7309 &stack_only(&text)[2..5],
7310 ["x64.fld_t [%3]", "x64.fchs", "x64.fstp_t [%4]"],
7311 "{text}"
7312 );
7313 }
7314
7315 #[test]
7316 fn comparing_two_long_doubles_puts_the_left_one_on_top() {
7317 let f64 = Type::float(rucc_ir::Float::F64);
7318 let (mut names, mut source, block, args) = blank(&[f64, f64]);
7319 let (left, right) = two_long_doubles(&mut source, block, &args);
7320 let mut build = Builder::new(&mut source, block);
7321 build.fcmp(FloatPred::Ogt, left, right, Flags::default());
7322 build.ret(&[]);
7323
7324 // `a > b`. `fucomip` asks about the top of the stack against what is under it, so the
7325 // operand the predicate is about has to go on last, which is the other way round from the
7326 // arithmetic above. The pop that clears the loser and the byte that reads the flags are
7327 // both inside the one opcode.
7328 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
7329 .expect("every instruction is written");
7330 let slots = pushed(&out);
7331 assert_eq!(slots, [2, 1], "the right operand goes on first and the left one on top");
7332 let text = mir::print_func(&out.func, &names, ®S);
7333 assert_eq!(
7334 &stack_only(&text)[4..],
7335 ["x64.fld_t [%6]", "x64.fld_t [%7]", "%8:gpr = x64.fucomip_set_a"],
7336 "{text}"
7337 );
7338 }
7339
7340 #[test]
7341 fn a_comparison_that_the_machine_has_backwards_swaps_the_two_pushes() {
7342 let f64 = Type::float(rucc_ir::Float::F64);
7343 let (mut names, mut source, block, args) = blank(&[f64, f64]);
7344 let (left, right) = two_long_doubles(&mut source, block, &args);
7345 let mut build = Builder::new(&mut source, block);
7346 build.fcmp(FloatPred::Olt, left, right, Flags::default());
7347 build.ret(&[]);
7348
7349 // `a < b` is `b > a` and this machine has the one condition, so the same opcode runs with
7350 // the operands the other way round. The same trade the vector rules make, and it has to
7351 // be the same one: a `long double` comparison that picked a different condition from the
7352 // `double` comparison of the same two numbers would be wrong at exactly the unordered
7353 // cases the two conditions differ on.
7354 //
7355 // Which slot each push names is the whole of the difference from the test above, and the
7356 // text does not show it, since an address in a frame is a `lea` with nothing in it until
7357 // `finish` has the numbers. So the slots are what is read here.
7358 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
7359 .expect("every instruction is written");
7360 let slots = pushed(&out);
7361 assert_eq!(slots, [1, 2], "the left operand goes on first and the right one on top");
7362 let text = mir::print_func(&out.func, &names, ®S);
7363 assert_eq!(
7364 &stack_only(&text)[4..],
7365 ["x64.fld_t [%6]", "x64.fld_t [%7]", "%8:gpr = x64.fucomip_set_a"],
7366 "{text}"
7367 );
7368 }
7369
7370 #[test]
7371 fn an_ordered_equal_needs_a_second_byte_to_put_the_two_conditions_together() {
7372 let f64 = Type::float(rucc_ir::Float::F64);
7373 let (mut names, mut source, block, args) = blank(&[f64, f64]);
7374 let (left, right) = two_long_doubles(&mut source, block, &args);
7375 let mut build = Builder::new(&mut source, block);
7376 build.fcmp(FloatPred::Oeq, left, right, Flags::default());
7377 build.ret(&[]);
7378
7379 // Equal and ordered are two conditions and the flags carry both, so the opcode writes a
7380 // second register as well as the one the value is in and ANDs them together. Said here by
7381 // handing it a spare, since an instruction that wrote a register nothing knew about would
7382 // be an instruction the allocator could put a live value in the way of.
7383 let text = lower(&mut names, &source);
7384 assert!(text.contains("%8:gpr, %9:gpr = x64.fucomip_set_e_and_np"), "{text}");
7385 }
7386
7387 #[test]
7388 fn a_comparison_that_is_never_asked_is_reported() {
7389 let f64 = Type::float(rucc_ir::Float::F64);
7390 let (mut names, mut source, block, args) = blank(&[f64, f64]);
7391 let (left, right) = two_long_doubles(&mut source, block, &args);
7392 let mut build = Builder::new(&mut source, block);
7393 build.fcmp(FloatPred::False, left, right, Flags::default());
7394 build.ret(&[]);
7395
7396 // Always false is a constant and not a comparison, so there is no condition to pick and
7397 // nothing here folds it into one: an instruction that quietly agreed with it would hide
7398 // that the optimizer left a comparison in that it should have taken out.
7399 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
7400 .expect_err("no condition is always false");
7401 assert_eq!(failed.to_string(), "no rule lowers a `fcmp` producing a `i1`");
7402 }
7403
7404 #[test]
7405 fn a_long_double_constant_is_the_bits_of_it_put_where_the_value_lives() {
7406 let (mut names, mut source, block, args) = blank(&[Type::PTR]);
7407 let mut build = Builder::new(&mut source, block);
7408 // `1.5L`, which is the leading bit and one more of significand, and an exponent of zero.
7409 let one_and_a_half = build.fconst(long_double(), 0x3fff_c000_0000_0000_0000);
7410 build.store(one_and_a_half, args[0], plain(), Flags::default());
7411 build.ret(&[]);
7412
7413 // No x87 instruction at all. A slot holding one of these is the value, so a constant is
7414 // its ten bytes written where the value lives, and whatever reads it does the `fld`.
7415 let text = lower(&mut names, &source);
7416 assert!(text.contains("x64.mov_ri_64 -4611686018427387904"), "{text}");
7417 assert!(text.contains("x64.mov_ri_16 16383"), "{text}");
7418 assert!(text.contains("x64.mov_mr_16 %3, [%1 + 8]"), "{text}");
7419 // The six bytes above the ten are the padding that makes the type sixteen wide, and they
7420 // are unspecified rather than zero, so nothing writes them.
7421 assert_eq!(text.matches("x64.mov_mr").count(), 2, "{text}");
7422 }
7423
7424 #[test]
7425 fn a_negative_long_double_constant_keeps_the_bit_above_its_exponent() {
7426 let (mut names, mut source, block, args) = blank(&[Type::PTR]);
7427 let mut build = Builder::new(&mut source, block);
7428 let minus = build.fconst(long_double(), 0xbfff_c000_0000_0000_0000);
7429 build.store(minus, args[0], plain(), Flags::default());
7430 build.ret(&[]);
7431
7432 // `-1.5L`. The sign is the top bit of the two byte half, so the immediate that half is put
7433 // in a register with is above the signed range of sixteen bits and has to stay there: read
7434 // as a number it would be negative, and it is not a number, it is two bytes.
7435 let text = lower(&mut names, &source);
7436 assert!(text.contains("x64.mov_ri_16 49151"), "{text}");
7437 }
7438
7439 #[test]
7440 fn a_long_double_crosses_an_edge_as_an_address_and_is_copied_where_it_lands() {
7441 let (mut names, mut source, block, args) = blank(&[Type::float(rucc_ir::Float::F64)]);
7442 let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
7443 let next = source.create_block();
7444 let param = source.append_param(next, long_double());
7445 Builder::new(&mut source, block).jump(next, &[wide]);
7446 Builder::new(&mut source, next).ret(&[param]);
7447
7448 // What the edge carries is the address of the slot the value is already in, which is an
7449 // ordinary register the allocator has an opinion about. The block on the other side copies
7450 // the sixteen bytes into a slot of its own before anything reads them, so a second edge
7451 // handing over a second address would still leave one place for a reader to look.
7452 let text = lower(&mut names, &source);
7453 let second: Vec<&str> = text
7454 .lines()
7455 .skip_while(|line| !line.starts_with("block1"))
7456 .skip(1)
7457 .take(3)
7458 .map(str::trim)
7459 .collect();
7460 assert_eq!(
7461 second,
7462 ["x64.fld_t [%4]", "%5:gpr = x64.lea_64 [$rsp]", "x64.fstp_t [%5]"],
7463 "{text}"
7464 );
7465 }
7466
7467 #[test]
7468 fn more_long_doubles_at_a_block_than_the_stack_is_deep_are_reported() {
7469 let f64 = Type::float(rucc_ir::Float::F64);
7470 let (mut names, mut source, block, args) = blank(&[f64]);
7471 let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
7472 let next = source.create_block();
7473 let params: Vec<Value> =
7474 (0..=X87_DEPTH).map(|_| source.append_param(next, long_double())).collect();
7475 let carried: Vec<Value> = params.iter().map(|_| wide).collect();
7476 Builder::new(&mut source, block).jump(next, &carried);
7477 Builder::new(&mut source, next).ret(&[params[0]]);
7478
7479 // The copies go through the x87 stack so that every one of them is read before any of them
7480 // is written, which is what makes a block that swaps two of these right. Nine of them do
7481 // not fit on the stack, and copying the ninth before or after the rest is the order that
7482 // could be wrong, so it is refused instead.
7483 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
7484 .expect_err("nine do not fit on the stack");
7485 assert_eq!(
7486 failed.to_string(),
7487 "block1 takes 9 parameters of type `f80` and only 8 can cross an edge at once"
7488 );
7489 assert_eq!(failed.inst(), None);
7490 }
7491}