Skip to main content

rucc_codegen/
lower.rs

1//! The selector: an IR function becomes a machine IR function.
2//!
3//! Design: `spec/10-backend.md` sections 10.2 and 10.3.
4//!
5//! What the matcher in [`crate::select`] does is answer one question about one term. What this
6//! does is ask it: walk a function, decide which terms are worth asking about, and build machine
7//! instructions out of what comes back. Nothing here decides what an IR term lowers to. That is
8//! in `rules/x86-64.rules` and it is proved before it is used, which is the whole point of the
9//! arrangement and the reason this file is short.
10//!
11//! # What it does with an instruction
12//!
13//! It tries the ways the instruction can be shown to the matcher, in order, and takes the first
14//! that a rule fires on. [`crate::term`] is what a way of showing one is, and the order is the
15//! most specific first: an operand that is a constant is offered as a constant before it is
16//! offered as a register, and an operand computed by an instruction of its own is offered as
17//! that instruction before it is offered as a register. A rule that wants an immediate too wide
18//! for the machine has a guard that turns it down, and the search carries on to the way of
19//! showing it that puts the constant in a register, which is the right answer and is one nobody
20//! had to write down.
21//!
22//! A constant is not lowered where it is written. It is materialized where a register for it is
23//! first wanted, which is what keeps a constant that every use folded into an immediate from
24//! leaving a dead instruction behind, and it also gives the value the shortest live range it
25//! could have. The instruction that materializes it comes from the rule set like everything else.
26//!
27//! # What it does not do yet
28//!
29//! Everything is in the general purpose registers, because every rule in the set is about an
30//! integer, so a call that passes a `double` and a function that returns one are both reported
31//! rather than lowered. So is an argument that travels on the stack, on either side of a call,
32//! and so is a call through an address rather than to a name.
33//!
34//! # A call
35//!
36//! Not a rule, because a rule pattern sees one term and what a call's operands are is whatever
37//! the signature made them. [`crate::abi`] builds one instead, out of the same description of the
38//! convention the arguments come from: the values it passes are reads constrained to the
39//! registers the convention places them in, what comes back is a write constrained to the
40//! register it comes back in, and every other register the callee is free to destroy is a write
41//! of that register and nothing else, which is all the allocator needs to keep a value out of it.
42//!
43//! What that costs the frame is an argument area, and nothing after selection could work out how
44//! big, so the size of the widest call is given back with the function. A function that makes no
45//! call at all is a leaf, and a leaf is the function that may use the red zone.
46//!
47//! # Where a block goes
48//!
49//! On the block, which is what machine IR does with an edge and is why the branches need no more
50//! rule language than the arithmetic did. A rule never names a block, so an unconditional jump
51//! has no rule at all and a conditional branch has one that is about its condition and nothing
52//! else. The arms are copied across after the block is filled, arguments and all, because an
53//! argument that is a constant is materialized where a register for it is first wanted and the
54//! end of the block is where an edge wants it.
55//!
56//! What this leaves behind is a function whose blocks are in the order the IR held them and whose
57//! branches are still branches on a register. Turning one into a `test` and a `jcc` is the block
58//! layout's, since which of the two arms falls through is the layout's answer, and [`crate::split`]
59//! has to run before allocation so that every edge carrying a value has somewhere to put it.
60//!
61//! A store and a return are the two things here that write no register. A store is emitted like
62//! everything else and the only difference is that there is no result to put anywhere, so the
63//! operands the target describes are all reads. A return is the same, and what it is for is its
64//! one operand: the target constrains it to the register the caller reads the value out of, and
65//! the allocator is what gets it there. The instruction that leaves is not chosen here at all,
66//! because the epilogue has to give the frame back first and [`crate::finish`] writes that after
67//! allocation, so a return of nothing is lowered to nothing.
68//!
69//! The entry block is the one block whose parameters are not block parameters here. They are the
70//! function's arguments, they are already somewhere when it starts, and [`crate::abi`] is what
71//! says where. An argument that arrives on the stack is reported rather than read, because where
72//! the stack put it is a distance into a frame and no frame exists until after allocation.
73//!
74//! Blocks are walked in the order the function holds them and a value is expected to be defined
75//! before it is used, which is true of the IR this is given because every pass before it keeps
76//! definitions ahead of uses.
77
78use std::collections::HashSet;
79use std::fmt;
80
81use rucc_base::{Interner, Symbol};
82use rucc_diag::Span;
83use rucc_ir::{
84    Abi, AsmOperand, AsmOperands, Block, Def, Extra, FloatPred, Func, Inst, Linkage, MemOrder,
85    Opcode, Param, PrefetchHint, RmwOp, Type, Value, Visibility,
86};
87use rucc_mir as mir;
88use rucc_target::x86_64;
89use rucc_target::{CallRegs, Constraint, OperandDesc, RegClass, Role, Segment};
90
91use crate::abi::{self, Missing, Refused};
92use crate::coverage::Fired;
93use crate::elsewhere::Elsewhere;
94use crate::frame::{Layout, Local};
95use crate::select::{Match, Piece, Rule, Table};
96use crate::term::{MAX_ARGS, PLAIN, Plan, Shown, Term, Terms};
97use crate::varargs;
98
99/// The prefix a rule file puts in front of a machine term, which says which target it belongs
100/// to and is not part of the opcode.
101pub(crate) const PREFIX: &str = "x64.";
102
103/// The instruction a global offset table slot is read with.
104///
105/// Not in [`x86_64::FRAME`] with the other opcodes this file names, because a frame has no use for
106/// it. It is spelled out here because the relocation it takes is only legal on a `mov` with a REX
107/// prefix, so the width is part of the requirement rather than a choice.
108const GOT_LOAD: &str = "mov_rm_64";
109
110/// How wide an address is on this target, which is the width a cast between a pointer and an
111/// integer has to be at for the cast to be nothing.
112const ADDRESS_BITS: u32 = 64;
113
114/// How many bytes a `long double` takes in memory, and what it is aligned to, which are the same
115/// number and are both more than the ten bytes that mean anything.
116///
117/// The psABI's answer rather than a choice here. `sizeof (long double)` is sixteen on this
118/// machine, so an array of them is laid out this way whatever a slot holding one does, and a slot
119/// that agreed with the array is one fewer thing to get wrong.
120const X87_BYTES: u32 = 16;
121
122/// How many values the x87 stack holds at once.
123///
124/// Eight, which is the machine's number rather than a choice here, and it matters in one place:
125/// the parameters of a block are copied through the stack so that they all move at once, and a
126/// block with more of them than this has nowhere to put the ninth.
127const X87_DEPTH: usize = 8;
128
129/// How many bytes a value passes through on its way between a register and the x87 stack.
130///
131/// Eight, because the widest thing that crosses is a `double` or a sixty four bit integer, and
132/// nothing crosses at eighty bits: a value that wide is already in the frame and the stack reaches
133/// it where it is.
134const X87_CROSSING: u32 = 8;
135
136/// Where the rounding field of the x87 control word is and what it has to be set to for the unit
137/// to cut towards zero, which is the one rounding C asks for that the unit does not do by default.
138///
139/// Both bits on is truncate. The field is ORed into the word that was already there rather than
140/// written over it, so the precision control and the exception masks somebody else set stay set.
141const X87_TRUNCATE: i64 = 0x0c00;
142
143/// Whether a type is the one this machine has no register for.
144///
145/// Only the eighty bit float is, and that is a fact about x86-64 rather than about floats: every
146/// other scalar the front end produces is in a general purpose register or a vector one, and this
147/// one is on the x87 stack while it is being worked on and in memory the rest of the time. So it
148/// has no place in [`Lowering::class_of`] and no name in [`crate::term`], and every instruction
149/// that touches one is written out by hand in this file.
150fn on_x87(ty: Type) -> bool {
151    ty.is_scalar() && ty.is_float() && ty.bits() == 80
152}
153
154/// Why a function could not be lowered.
155///
156/// One reason and then nothing. A function with no rule for something in it is a function this
157/// cannot finish, and the second thing it could not lower is not news.
158#[derive(Debug, Clone, PartialEq, Eq)]
159pub enum Unsupported {
160    /// An instruction no rule fires on.
161    Inst {
162        /// The instruction that stopped it.
163        inst: Inst,
164        /// What the rule file would call it, or nothing if the rule language has no name for it
165        /// at all, which is what an instruction at a width nothing is written about looks like.
166        term: Option<&'static str>,
167        /// The opcode, which is what gets named when the rule language has no word for it.
168        ///
169        /// An opcode the rule language has no word for is exactly the opcode no rule lowers, so
170        /// without this the message would be empty in every case where somebody needs it.
171        opcode: Opcode,
172        /// What it produces, or nothing for an instruction that is only an effect.
173        ty: Option<Type>,
174    },
175    /// A parameter that does not arrive somewhere this can bring it in from.
176    ///
177    /// Not an instruction, which is why it is a separate arm: it is a fact about the signature
178    /// and there is nothing in the body of the function to point at.
179    Argument {
180        /// Its position in the signature.
181        index: usize,
182        /// What is wrong with where it arrives.
183        missing: Missing,
184    },
185    /// A call that passes or gives back a value this cannot put where the convention wants it.
186    Call {
187        /// The call.
188        inst: Inst,
189        /// Which value, and what is wrong with where it travels.
190        refused: Refused,
191    },
192    /// A `return` this cannot put where the convention wants it.
193    ///
194    /// A separate arm from [`Unsupported::Inst`] because it is not an instruction no rule fires
195    /// on. A return of more than one value is built from the convention rather than matched, the
196    /// same way a call is, so what goes wrong with one is what goes wrong with a call and not the
197    /// absence of a rule.
198    Returned {
199        /// The `return`.
200        inst: Inst,
201        /// What is wrong with where one of the values travels.
202        missing: Missing,
203    },
204    /// A stack slot the frame cannot give the bytes it asked for.
205    ///
206    /// Not an instruction no rule covers. An `alloca` is built here rather than matched, so what
207    /// goes wrong with one is what the frame can and cannot hold rather than what the rules spell.
208    Dynamic {
209        /// The `alloca`.
210        inst: Inst,
211        /// What the frame could not do about it.
212        growing: Growing,
213    },
214    /// More parameters of a type that travels on the x87 stack than the stack is deep.
215    ///
216    /// Not an instruction either, for the reason a function's parameter is not one: it is a fact
217    /// about the block and there is nothing in the block to point at. What crosses an edge for one
218    /// of these is the address of where the value is, and the block copies the bytes into a slot
219    /// of its own, all of them through the stack at once so that a block carrying two of them
220    /// swapped is copied in an order that is right. Eight is as many as the stack holds, and a
221    /// ninth would have to be copied before or after the rest, which is the order that could be
222    /// wrong.
223    Phi {
224        /// Which block it arrives at.
225        block: Block,
226        /// How many of them arrive there, which is the whole of what is wrong.
227        count: usize,
228        /// What they are.
229        ty: Type,
230    },
231    /// An `asm` statement this cannot build.
232    ///
233    /// Not an instruction no rule fires on, for the reason a call is not one: what it stands for is
234    /// whatever its template says, and no pattern over terms can read a string.
235    Assembly {
236        /// The `inline_asm`.
237        inst: Inst,
238        /// What about it is not built here yet.
239        refused: Written,
240    },
241}
242
243/// What about an `asm` statement is not built yet.
244#[derive(Debug, Clone, Copy, PartialEq, Eq)]
245pub enum Written {
246    /// A template with instructions in it.
247    Template,
248    /// An `asm goto`, whose labels make the statement a terminator.
249    Goto,
250    /// An operand this cannot put where the constraint says it goes.
251    Operand,
252}
253
254impl Written {
255    /// The rest of the sentence that starts with the statement.
256    #[must_use]
257    pub fn why(self) -> &'static str {
258        match self {
259            // The template is the assembler's to read and there is no assembler here yet, so a
260            // template with anything in it is a string nothing can turn into bytes. An empty one is
261            // no instructions, and no instructions is something this can write.
262            Written::Template => "has instructions in its template, which nothing here assembles",
263            Written::Goto => "jumps to a label, which nothing here builds an edge for",
264            Written::Operand => "has an operand this cannot place",
265        }
266    }
267}
268
269/// What the frame could not do about a stack slot.
270#[derive(Debug, Clone, Copy, PartialEq, Eq)]
271pub enum Growing {
272    /// An object of a size the number a frame counts bytes in does not reach.
273    Huge,
274    /// A variable length array wanting more alignment than a call leaves the stack pointer with.
275    ///
276    /// Rounding the stack pointer down again after the bytes have been taken would put it
277    /// somewhere no constant reaches the rest of the frame from, so a frame like this needs a
278    /// second base register held for the whole of the function. Nothing here holds one.
279    Aligned,
280    /// A variable length array in a function whose frame is meant to be touched a page at a time.
281    ///
282    /// The pages a prologue takes are touched by the prologue, which knows how many there are when
283    /// it is written. The pages a variable length array takes are not known until the declaration
284    /// runs, so touching them is a loop next to the declaration, and there is no loop here yet.
285    Probed,
286}
287
288impl Growing {
289    /// The rest of the sentence that starts with the slot.
290    #[must_use]
291    pub fn why(self) -> &'static str {
292        match self {
293            Growing::Huge => "is more bytes than a frame counts",
294            Growing::Aligned => {
295                "wants more alignment than the stack pointer is left on, which needs a base \
296                 register nothing here keeps"
297            }
298            Growing::Probed => {
299                "grows the stack, and nothing here touches the pages it takes a page at a time"
300            }
301        }
302    }
303}
304
305impl Unsupported {
306    /// The instruction it is about, or nothing for the one arm that is about a signature.
307    ///
308    /// What a caller wants this for is the span. The function knows where every instruction in
309    /// it came from, so a caller holding both can point a message at the line somebody wrote
310    /// rather than at the file as a whole, and nothing here has to carry a span of its own.
311    pub fn inst(&self) -> Option<Inst> {
312        match *self {
313            Unsupported::Inst { inst, .. }
314            | Unsupported::Call { inst, .. }
315            | Unsupported::Returned { inst, .. }
316            | Unsupported::Dynamic { inst, .. }
317            | Unsupported::Assembly { inst, .. } => Some(inst),
318            Unsupported::Argument { .. } | Unsupported::Phi { .. } => None,
319        }
320    }
321}
322
323impl fmt::Display for Unsupported {
324    fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
325        match *self {
326            Unsupported::Inst { term: Some(term), .. } => write!(f, "no rule lowers `{term}`"),
327            Unsupported::Inst { term: None, opcode, ty: Some(ty), .. } => {
328                write!(f, "no rule lowers a `{opcode}` producing a `{ty}`")
329            }
330            Unsupported::Inst { term: None, opcode, ty: None, .. } => {
331                write!(f, "no rule lowers a `{opcode}`")
332            }
333            Unsupported::Argument { index, missing } => {
334                write!(f, "parameter {index} {}", missing.why())
335            }
336            Unsupported::Call { refused: Refused { argument: Some(index), missing }, .. } => {
337                write!(f, "argument {index} of this call {}", missing.why())
338            }
339            Unsupported::Call { refused: Refused { argument: None, missing }, .. } => {
340                write!(f, "what this call gives back {}", missing.why())
341            }
342            Unsupported::Returned { missing, .. } => {
343                write!(f, "what this function gives back {}", missing.why())
344            }
345            Unsupported::Dynamic { growing, .. } => {
346                write!(f, "this local {}", growing.why())
347            }
348            Unsupported::Phi { block, count, ty } => {
349                let block = block.index();
350                write!(
351                    f,
352                    "block{block} takes {count} parameters of type `{ty}` and only {X87_DEPTH} can cross an edge at once"
353                )
354            }
355            Unsupported::Assembly { refused, .. } => write!(f, "this `asm` {}", refused.why()),
356        }
357    }
358}
359
360impl std::error::Error for Unsupported {}
361
362/// A lowered function, and what the frame needs that the machine IR does not hold.
363#[derive(Debug)]
364pub struct Lowered {
365    /// The function, in machine instructions.
366    pub func: mir::Func,
367    /// What it wants its stack to look like, which is separate from the function so that the two
368    /// can be read and written at the same time.
369    pub stack: Stack,
370    /// Which rules of the table lowered it, which is what `-Zrule-coverage` asks for and what
371    /// `crate::coverage` writes down.
372    pub fired: Fired,
373    /// Which machine IR block each IR block became, indexed by the IR block's own index, and
374    /// nothing for a block the walk never reached.
375    ///
376    /// Here because it is the only place the correspondence exists. Selection makes one block per
377    /// block, in the same order and with the arms in the same order, so anything the IR knows
378    /// about a block can be carried down through this and nothing else, and
379    /// [`crate::weights::carry`] is what does.
380    pub blocks: Vec<Option<mir::Block>>,
381}
382
383/// What a function's stack has to hold, as far as selection is able to say.
384///
385/// All of it is answered here because selection is where a call is built and where an `alloca`
386/// is read, and nothing after it could tell what either of them needed.
387#[derive(Debug, Default)]
388pub struct Stack {
389    /// How many bytes the widest call in the function needs below the stack pointer for the
390    /// arguments it passes there, or `None` for a function that makes no call at all.
391    ///
392    /// `None` is a leaf, which is the function that may use the red zone and the one whose stack
393    /// pointer does not have to be left aligned for anybody.
394    pub calls: Option<u32>,
395    /// The memory the function asked for itself, one entry for every `alloca` in it, in the order
396    /// the walk reached them.
397    pub locals: Vec<Local>,
398    /// Which instruction computes the address of which of those locals.
399    ///
400    /// An address in the frame is a distance from the stack pointer, and there is no frame until
401    /// after allocation, so the instruction is written here with nothing in its displacement and
402    /// [`crate::finish`] writes the number in once [`crate::frame::Frame`] knows it.
403    pub addresses: Vec<(mir::Inst, usize)>,
404    /// Which instruction computes the address of a piece of memory whose size the function works
405    /// out while it runs, which is what a variable length array is.
406    ///
407    /// Waiting on [`crate::finish`] for a different number from the one the addresses above are:
408    /// the bytes were taken off the stack pointer by the instruction in front of this one, so where
409    /// they start is however much of the bottom of the frame belongs to the arguments of a call,
410    /// and that is not known until the frame is.
411    pub dynamic: Vec<mir::Inst>,
412    /// Where the function first moves the stack pointer while it runs, if it does at all.
413    ///
414    /// Two things are read off this. One is whether at all, which is what [`crate::frame::Layout`]
415    /// wants, because a frame that moves its stack pointer has a different shape from one that does
416    /// not and the layout is built before the instructions are looked at again. See `Growing` in
417    /// [`crate::frame`]. The other is where, so that a caller that cannot accept such a frame has
418    /// somewhere to point when it says so.
419    pub grown_at: Option<Inst>,
420    /// Which instruction reads which of the arguments the caller passed on the stack, as how far up
421    /// the caller's argument area it reads.
422    ///
423    /// Waiting on [`crate::finish`] for the same reason the addresses above are, and on one thing
424    /// more: where the caller's argument area is from inside this function depends on whether the
425    /// prologue had to force the stack pointer's alignment, so which register the load reads
426    /// through is not settled here either.
427    pub arguments: Vec<(mir::Inst, u32)>,
428}
429
430impl Stack {
431    /// The layout given, with the three fields only the lowering knows the answer to filled in.
432    ///
433    /// Everything else in a layout comes from the flags the function is compiled under or from the
434    /// allocation, so this takes one and returns it rather than building one.
435    #[must_use]
436    pub fn layout<'a>(&'a self, base: Layout<'a>) -> Layout<'a> {
437        Layout {
438            leaf: self.calls.is_none(),
439            outgoing: self.calls.unwrap_or(0),
440            locals: &self.locals,
441            grows: self.grown_at.is_some(),
442            ..base
443        }
444    }
445}
446
447/// The x86-64 machine IR for that function.
448///
449/// # Errors
450///
451/// The first instruction no rule fires on, which today is anything at a width the rule set is not
452/// written at, a parameter that does not arrive in a register this can read, or a call that
453/// passes something this cannot put where the convention wants it.
454pub fn func(
455    source: &Func,
456    names: &mut Interner,
457    conv: &'static CallRegs,
458    elsewhere: &Elsewhere,
459) -> Result<Lowered, Unsupported> {
460    Lowering::new(source, names, conv, elsewhere).run()
461}
462
463/// What the matcher settled on for one block, indexed the way the block's instructions are.
464struct Decided {
465    /// What each instruction matched, and nothing for one that matched no rule or was folded
466    /// into a later one.
467    found: Vec<Option<Match<Term>>>,
468    /// How each instruction showed its operands to the matcher, which is what says what it took.
469    plans: Vec<Option<Plan>>,
470    /// The instructions some other instruction took, which are the ones with nothing to write.
471    folded: Vec<Inst>,
472}
473
474/// One function being lowered.
475struct Lowering<'a> {
476    source: &'a Func,
477    names: &'a mut Interner,
478    out: mir::Func,
479    /// The machine register each IR value is in, once it has one.
480    regs: Vec<Option<mir::Reg>>,
481    /// For a constant that has been written into a register, the block it was written into,
482    /// which is the only block that register is any good in.
483    written: Vec<Option<mir::Block>>,
484    /// How many times each IR value is read, which is what says whether an instruction may be
485    /// folded into the one that reads it.
486    uses: Vec<u32>,
487    /// The block being filled.
488    at: Option<mir::Block>,
489    /// The machine IR block each IR block became.
490    blocks: Vec<Option<mir::Block>>,
491    /// The class an address is in, which is the general purpose one and is not a question: every
492    /// register an addressing mode names holds part of an address, and there is no machine here
493    /// that computes an address anywhere but in this file. Which class a *value* is in is
494    /// [`Lowering::class_of`], and it is a question, because a float is in the other one.
495    gpr: RegClass,
496    /// Where the convention this function is compiled for puts things, which is read for the
497    /// arguments and for the calls.
498    conv: &'static CallRegs,
499    /// Which names this function may not work an address out for itself, which is a fact about the
500    /// module and so is worked out before any of this and handed in.
501    elsewhere: &'a Elsewhere,
502    /// What the function wants its stack to look like, filled in as the walk finds out.
503    stack: Stack,
504    /// What a `va_start` in this function has to write, or nothing for a function that takes no
505    /// arguments its signature does not name.
506    ///
507    /// Worked out once, when the entry block binds the parameters, because every number in it is
508    /// about where those parameters left the walk over the argument registers and there is nowhere
509    /// else that knows.
510    varargs: Option<Varargs>,
511    /// Which of the function's stack objects each eighty bit value lives in, once it has asked
512    /// for one.
513    ///
514    /// One slot per value and it is never given back, which is what makes an eighty bit value
515    /// behave like every other one: it is written once and read wherever it is read, and no two
516    /// of them share a slot the way two of them would share a register. What is in a register is
517    /// the address, and that is worked out again at every use rather than kept, so nothing here
518    /// holds a general purpose register open across a whole function.
519    slots: Vec<Option<usize>>,
520    /// The eight bytes a value passes through between a register and the x87 stack, once
521    /// something has wanted them.
522    ///
523    /// One for the whole function, because every group that uses it is a handful of instructions
524    /// with nothing in between: the bytes are written, read straight back and never looked at
525    /// again, so a second slot would be a second slot holding the same nothing.
526    crossing: Option<usize>,
527    /// The four bytes the control word is saved in and the changed copy written to, once
528    /// something has wanted them.
529    ///
530    /// One for the whole function for the reason above, and four rather than two because it is
531    /// two words: the one the unit had and the one with the rounding field turned to truncate.
532    control: Option<usize>,
533    /// Which rules have fired so far.
534    fired: Fired,
535}
536
537/// What a `va_start` in a variadic function writes into the list it is given.
538///
539/// Three of the four are settled here and the fourth is not a number at all yet: where the save
540/// area is and where the caller's argument area is are both distances into a frame that does not
541/// exist until after allocation, so both are `lea` instructions [`crate::finish`] fills in.
542#[derive(Debug, Clone, Copy, PartialEq, Eq)]
543struct Varargs {
544    /// Which of the function's stack objects is the register save area.
545    save: usize,
546    /// How far up the caller's argument area the first argument the signature does not name is,
547    /// which is the whole of that area the named ones did not take.
548    incoming: u32,
549    /// What `gp_offset` starts at, which is past the general purpose registers the named arguments
550    /// took.
551    integers: u32,
552    /// What `fp_offset` starts at, which is past the vector ones.
553    floats: u32,
554}
555
556/// How far a function's name reaches, narrowed from the linkage the IR gave it.
557///
558/// The IR has five and an object file says three, and the two the linker cannot tell apart are
559/// the two weak ones: which of them a symbol had is a fact the optimizer reads and the linker has
560/// no way to record. A function is never `Common`, since that is what a tentative definition of an
561/// object is and there is no tentative definition of a function, and it is written here rather
562/// than left out so that a linkage added later has to come past this.
563const fn binding(linkage: Linkage) -> mir::Binding {
564    match linkage {
565        Linkage::Internal => mir::Binding::Local,
566        Linkage::Weak | Linkage::LinkOnce => mir::Binding::Weak,
567        Linkage::External | Linkage::Common => mir::Binding::Global,
568    }
569}
570
571/// How far a function's name reaches outside a shared library, carried across unchanged.
572///
573/// Nothing is narrowed here the way [`binding`] narrows the linkage, because ELF records all
574/// three of these and the two enumerations are the same three answers written twice: once in a
575/// crate that is not allowed to know what an object file is and once in one that is.
576const fn visibility(visibility: Visibility) -> mir::Visibility {
577    match visibility {
578        Visibility::Default => mir::Visibility::Default,
579        Visibility::Hidden => mir::Visibility::Hidden,
580        Visibility::Protected => mir::Visibility::Protected,
581    }
582}
583
584impl<'a> Lowering<'a> {
585    fn new(
586        source: &'a Func,
587        names: &'a mut Interner,
588        conv: &'static CallRegs,
589        elsewhere: &'a Elsewhere,
590    ) -> Self {
591        let counts = source.counts();
592        let name = source.name;
593        let mut uses = vec![0; counts.values];
594        for block in source.blocks() {
595            for inst in source.insts(block) {
596                for &arg in &source[source[inst].args] {
597                    uses[arg.index()] += 1;
598                }
599                for call in source.successors(inst) {
600                    for &arg in &source[call.args] {
601                        uses[arg.index()] += 1;
602                    }
603                }
604            }
605        }
606        let mut out = mir::Func::new(name);
607        out.align = source.align;
608        out.binding = binding(source.linkage);
609        out.visibility = visibility(source.visibility);
610        Self {
611            source,
612            names,
613            out,
614            regs: vec![None; counts.values],
615            written: vec![None; counts.values],
616            blocks: vec![None; counts.blocks],
617            uses,
618            at: None,
619            gpr: x86_64::GPR,
620            conv,
621            elsewhere,
622            stack: Stack::default(),
623            varargs: None,
624            slots: vec![None; counts.values],
625            crossing: None,
626            control: None,
627            fired: Fired::new(),
628        }
629    }
630
631    fn run(mut self) -> Result<Lowered, Unsupported> {
632        // Every block before any of them is filled, because a block that jumps forward has to
633        // name the block it jumps to and a machine IR block is named by a handle rather than by
634        // the IR block it came from.
635        for block in self.source.blocks() {
636            let out = self.out.create_block();
637            self.blocks[block.index()] = Some(out);
638        }
639        for block in self.order() {
640            self.block(block)?;
641        }
642        // And the name each block an image holds the address of was given, which nothing in the
643        // walk above would ask for: the `lea` a label address is inside the function needs no
644        // symbol, and the one thing that does is a relocation in another section.
645        let named: Vec<(Block, Symbol)> = self.source.named_blocks().collect();
646        let labels: Vec<(mir::Block, Symbol)> =
647            named.into_iter().map(|(block, name)| (self.out_block(block), name)).collect();
648        self.out.labels = labels;
649        Ok(Lowered { func: self.out, stack: self.stack, fired: self.fired, blocks: self.blocks })
650    }
651
652    /// The order the blocks are filled in, which is not the order they are written in.
653    ///
654    /// Reverse postorder, because a value is written in a block that dominates every block that
655    /// reads it and a block in reverse postorder comes before every block it dominates. The order
656    /// the blocks are written in does not have that property: a block written early can read a
657    /// value a block below it writes, and reading a value with no register yet mints one, so the
658    /// register the definition writes later is not the register the read named. Nothing writes the
659    /// one the read named, and what comes out is a function that loads a stack slot no store ever
660    /// reached. It is the order this walk goes in rather than the order the blocks come out in,
661    /// which is what the loop above fixes, so the machine function is still written the way the IR
662    /// function was.
663    ///
664    /// Blocks the entry does not reach come last, in the order they are written in. Nothing runs
665    /// them and nothing they name is read by anything that does, but they still have to be filled,
666    /// because a machine block with no terminator is not one the passes below can read.
667    fn order(&self) -> Vec<Block> {
668        let Some(entry) = self.source.entry() else { return self.source.blocks().collect() };
669        let count = self.blocks.len();
670        let mut succs: Vec<Vec<Block>> = vec![Vec::new(); count];
671        for block in self.source.blocks() {
672            let Some(term) = self.source.terminator(block) else { continue };
673            succs[block.index()] = self.source.successors(term).map(|call| call.block).collect();
674        }
675        // An explicit stack, because the depth of the walk is the number of blocks and a function
676        // built by a generator has as many of those as it likes.
677        let mut seen = vec![false; count];
678        let mut order = Vec::with_capacity(count);
679        let mut stack = vec![(entry, 0usize)];
680        seen[entry.index()] = true;
681        while let Some((block, at)) = stack.pop() {
682            let Some(&next) = succs[block.index()].get(at) else {
683                order.push(block);
684                continue;
685            };
686            stack.push((block, at + 1));
687            if !seen[next.index()] {
688                seen[next.index()] = true;
689                stack.push((next, 0));
690            }
691        }
692        order.reverse();
693        order.extend(self.source.blocks().filter(|block| !seen[block.index()]));
694        order
695    }
696
697    /// One block: its parameters, then every instruction in it that is not folded into another.
698    fn block(&mut self, block: Block) -> Result<(), Unsupported> {
699        let out = self.out_block(block);
700        self.at = Some(out);
701        if self.source.entry() == Some(block) {
702            self.arrive(block, out)?;
703        } else {
704            let mut arriving = Vec::new();
705            for &param in &self.source[block].params {
706                // A value with no register to arrive in, which the class would not say, since
707                // `class_of` puts one of these in the general purpose file on purpose and what it
708                // means by that is that nothing there can hold it. What crosses the edge for one
709                // of those is the address of where the value already is, so the parameter is a
710                // pointer here and the bytes it points at are copied below.
711                let ty = self.source[param].ty;
712                let reg = self.out.append_param(out, self.class_of(ty));
713                self.regs[param.index()] = Some(reg);
714                if on_x87(ty) {
715                    arriving.push((param, reg));
716                }
717            }
718            self.settle(block, &arriving)?;
719        }
720
721        // What each instruction matched, and which instructions were folded into another. The
722        // decision is made for the whole block before any of it is written, and it is made more
723        // than once: a value that only some of its readers took has to be put back in a register
724        // for all of them, and taking it away from those readers changes what they match.
725        let insts: Vec<Inst> = self.source.insts(block).collect();
726        let mut refused: HashSet<Value> = HashSet::new();
727        let mut decided = self.decide(&insts, &refused);
728        while let Some(value) = self.left_alive(&insts, &decided.plans) {
729            refused.insert(value);
730            decided = self.decide(&insts, &refused);
731        }
732        let Decided { found, folded, .. } = decided;
733
734        for (&inst, matched) in insts.iter().zip(found) {
735            if folded.contains(&inst) || self.writes_nothing(inst) {
736                continue;
737            }
738            // A call is built from the convention rather than matched, which is why it is the one
739            // opcode looked at by name here. Through an address it is a different instruction and
740            // the same convention, so the two arrive at the same place and differ in one line of
741            // it.
742            match self.source[inst].opcode {
743                Opcode::Call | Opcode::CallIndirect => {
744                    self.called(inst)?;
745                    continue;
746                }
747                // Built from the frame rather than matched, for the same shape of reason a call
748                // is built from the convention: what a rule replaces a term with is instructions,
749                // and what an `alloca` needs first is bytes, which the rule language has no way
750                // to ask for.
751                Opcode::Alloca => {
752                    self.reserve(inst)?;
753                    continue;
754                }
755                // Reading the stack pointer and writing it back, which are the two ends of a scope
756                // holding a variable length array. Built here for the reason an `alloca` is: the
757                // value is a register the rule language has no way to name, because what it holds
758                // is not a value the program computed but where the machine's stack had got to.
759                Opcode::StackSave => {
760                    self.stack_pointer(inst, false)?;
761                    continue;
762                }
763                Opcode::StackRestore => {
764                    self.stack_pointer(inst, true)?;
765                    continue;
766                }
767                // The address of a name, built here for the same reason an `alloca` is: what a
768                // rule replaces a term with is instructions over values, and the operand of this
769                // one is a symbol, which is a thing the rule language has no way to bind and the
770                // solver has no way to say anything about. There is nothing in `lea sym(%rip)` a
771                // proof over bitvectors could discharge, because what makes it the right answer
772                // is the relocation and what the linker does with it.
773                Opcode::GlobalAddr => {
774                    self.address_of(inst)?;
775                    continue;
776                }
777                // The address of a label and the branch that reads one, built here for the same
778                // reason and for one more. The reason is the same: what the first of them names is
779                // a block, which is not a value a rule pattern can bind, and there is nothing in
780                // the distance between two places in one function that a proof over bitvectors
781                // could discharge. The extra one is that the second is a terminator whose arms are
782                // not two and not fixed, and a rule says what an instruction reads rather than
783                // where a block goes.
784                Opcode::BlockAddr => {
785                    self.block_address(inst)?;
786                    continue;
787                }
788                Opcode::IndirectBr => {
789                    self.indirect_branch(inst)?;
790                    continue;
791                }
792                // Where this thread's own storage starts, built here for a reason of the same
793                // shape: what it reads is `%fs`, which is not a register the rule language can
794                // bind and not one a proof over bitvectors could say anything about, because what
795                // makes the load the right answer is an agreement between the loader and the C
796                // library rather than any arithmetic.
797                Opcode::ThreadPointer => {
798                    self.thread_pointer(inst)?;
799                    continue;
800                }
801                // Built from the frame for the reason an `alloca` is, and from the convention for
802                // the reason a call is: three of the four fields it writes are distances that do
803                // not exist until the frame does, and the fourth is where the walk over the
804                // argument registers stopped. A function that is not variadic has no such walk to
805                // report, so it has nothing here and is refused below, which is the right answer
806                // for a `va_start` in one.
807                Opcode::VaStart if self.varargs.is_some() => {
808                    self.va_start(inst)?;
809                    continue;
810                }
811                // A return of more than one value, which is a structure small enough to come
812                // back in a pair of registers. Built from the convention for the reason a call
813                // is: which register each half goes in depends on the halves in front of it,
814                // because the two register files are walked separately, and a pattern over a term
815                // cannot see them. A return of one value is a term with a name and a rule, and it
816                // stays one.
817                //
818                // A return of none in a function whose answer went through memory is here too,
819                // and for a different reason: what it gives back is not written in the IR at all.
820                // The convention says the address the caller handed over comes back, and only the
821                // signature says this function was handed one.
822                //
823                // And a return of one eighty bit value, for a third reason: what a rule would
824                // write is an instruction leaving the value in a register, and this one is left on
825                // the x87 stack instead. A rule could not name that stack any more than any other
826                // rule about this type could.
827                Opcode::Return
828                    if self.source[self.source[inst].args].len() > 1
829                        || self.sret().is_some()
830                        || self.gives_back_x87(inst) =>
831                {
832                    self.returned(inst)?;
833                    continue;
834                }
835                // A cast between a pointer and an integer of the same width, which on this
836                // machine is every one the front end writes. No instruction at all, so no rule
837                // could name one.
838                Opcode::PtrToInt | Opcode::IntToPtr => {
839                    self.rename(inst)?;
840                    continue;
841                }
842                // A barrier, which is one instruction or none depending on the ordering. Written
843                // by name because there is nothing about it a rule could be proved against, the
844                // way there is nothing to prove about the address of a symbol.
845                Opcode::Fence => {
846                    self.barrier(inst)?;
847                    continue;
848                }
849                // A hint, written by name for the reason a barrier is and one step further: not
850                // only is there no equality for a proof to discharge, there is nothing about the
851                // program around it either. Which of the four instructions it is comes out of the
852                // number the builtin was given, which is beside the instruction rather than in it.
853                Opcode::Prefetch => {
854                    self.hint(inst)?;
855                    continue;
856                }
857                // Stopping, written by name for the first half of the barrier's reason: it
858                // computes nothing, so there is no term for a rule to replace, and what makes it
859                // right is what the operating system does with the fault rather than anything a
860                // proof over bitvectors could discharge.
861                Opcode::Trap => {
862                    self.trap(inst);
863                    continue;
864                }
865                // A compare and exchange, which is written by name because it produces two values
866                // and a rule produces one. The replacement of a rule is one term, a term names the
867                // value an instruction computes, and there is no way in that language to say that
868                // an instruction leaves an answer in one place and a yes or no in another.
869                Opcode::Cmpxchg => {
870                    self.exchange(inst)?;
871                    continue;
872                }
873                // A read modify write, which is written by name for a different reason: it produces
874                // one value, so a rule could name it, and what it does is not in the head a rule
875                // matches on. Every one of the thirteen operations is the same opcode at the same
876                // type and differs only in what is carried beside it, so one pattern would be all
877                // thirteen patterns. Of the thirteen only the three with an instruction reach here,
878                // since `crate::retry` turned the rest into loops a long way above this.
879                Opcode::AtomicRmw => {
880                    self.modify(inst)?;
881                    continue;
882                }
883                // An `asm` statement, whose lowering is its template and there is no term for a
884                // string. Written by name for the reason a barrier is, and before the x87 arm
885                // below so that an `asm` holding a `long double` is refused as the `asm` it is
886                // rather than as an instruction nothing computes.
887                Opcode::InlineAsm => {
888                    self.assembly(inst)?;
889                    continue;
890                }
891                // Anything at all with an eighty bit float in it, which is the one arm here
892                // chosen by a type rather than by an opcode, because what makes these different
893                // is not what they do but where the value is. A `long double` has no register,
894                // so it has no name in `crate::term` and no rule could bind one: every one of
895                // these is a group of instructions over a frame slot, written out below.
896                //
897                // Last of the arms, so that a call and a return with one of these in them reach
898                // the convention first and are refused by it, which is the truer answer: what is
899                // wrong there is where the value has to travel and not that nothing can compute
900                // it.
901                _ if self.touches_x87(inst) => {
902                    self.x87(inst)?;
903                    continue;
904                }
905                _ => {}
906            }
907            let matched = matched.ok_or_else(|| self.unsupported(inst))?;
908            self.emit(inst, &matched)?;
909            // After it is built rather than when it matched, so that what is recorded is the rules
910            // this function was lowered by and not the rules something was tried with.
911            self.fired.mark(matched.rule);
912        }
913        self.edges(block, out)
914    }
915
916    /// One call, which is built from the convention rather than matched against the table for the
917    /// same reason the arguments of the function itself are.
918    ///
919    /// The arguments are read before the call is built, which is what materializes a constant
920    /// argument into a register, since no call passes an immediate.
921    ///
922    /// A call to a name and a call through an address are both here, and what tells them apart is
923    /// the opcode rather than whether a callee was recorded, which is the same thing the verifier
924    /// reads. Through an address the first operand is the address and the arguments are the ones
925    /// behind it, and everything after that is the same: where each argument goes, where the value
926    /// comes back and which registers are gone across it are the convention's answers and the
927    /// convention does not ask what is being called.
928    fn called(&mut self, inst: Inst) -> Result<(), Unsupported> {
929        let data = &self.source[inst];
930        let Extra::Call(info) = data.extra else { return Err(self.unsupported(inst)) };
931        let info = self.source[info];
932        let indirect = data.opcode == Opcode::CallIndirect;
933
934        let values: Vec<Value> = self.source[data.args].to_vec();
935        let callee = if indirect {
936            let &address = values.first().ok_or_else(|| self.unsupported(inst))?;
937            abi::Callee::Through(self.reg_of(address)?)
938        } else {
939            abi::Callee::Named(info.callee.ok_or_else(|| self.unsupported(inst))?)
940        };
941
942        // What the ABI asks of each argument, read out before any of them is, because reading one
943        // borrows the function this is a table in. The ones the signature names are the signature's
944        // answer and the ones behind them are the call's, which is where a structure passed to a
945        // variadic callee by value says that its bytes travel: there is no parameter to say it on.
946        let signature = &self.source[info.signature];
947        let variadic = signature.variadic;
948        let named: Vec<Abi> = signature.params.iter().map(|param| param.abi).collect();
949        let beyond: Vec<Abi> = self.source[info.varargs].to_vec();
950        // Every value that comes back and not only the first. A structure small enough to travel
951        // in registers comes back in up to two of them, and which register each half is in is the
952        // convention's answer, which is why the whole list goes to the same place the arguments do
953        // rather than to a rule.
954        let returns: Vec<Type> = signature.return_types().collect();
955
956        let mut args = Vec::with_capacity(values.len());
957        for (index, value) in values.into_iter().skip(usize::from(indirect)).enumerate() {
958            let abi = named.get(index).or_else(|| beyond.get(index - named.len()));
959            let abi = abi.copied().unwrap_or_default();
960            let ty = self.source[value].ty;
961            // What travels for an eighty bit value is its bytes, so what the call is handed is
962            // where they are rather than a register they are in, and there is no register they
963            // could be in. Everything else about it is a sixteen byte object passed by value and
964            // is built by the same code.
965            let reg =
966                if abi::on_the_stack(ty) { self.x87_slot(value) } else { self.reg_of(value)? };
967            args.push(abi::Passing { ty, reg, abi });
968        }
969        let block = self.at.expect("a block is being filled");
970        let what = abi::Calling { callee, args: &args, returns: &returns, variadic };
971        let made = abi::call(&mut self.out, block, &what, self.conv, self.names)
972            .map_err(|refused| Unsupported::Call { inst, refused })?;
973        let calls = &mut self.stack.calls;
974        *calls = Some(calls.unwrap_or(0).max(made.outgoing));
975        // An eighty bit value came back on the x87 stack, and the one thing that has to happen
976        // before anything else touches that stack is taking it off. So the `fstp` goes here, in
977        // front of everything the block does next, and after it the value is in its slot and is
978        // read the way every other one is.
979        let results: Vec<Value> = self.source[inst].results().collect();
980        if let [result] = results[..] {
981            if abi::on_the_stack(self.source[result].ty) {
982                let span = self.source.span(inst);
983                let into = self.x87_slot(result);
984                let into = self.through(into);
985                self.x87_at("fstp_t", span, into);
986                return Ok(());
987            }
988        }
989        for (result, &reg) in results.into_iter().zip(&made.results) {
990            self.regs[result.index()] = Some(reg);
991        }
992        Ok(())
993    }
994
995    /// The pointer a function returning through memory was handed, or nothing in a function that
996    /// was not.
997    ///
998    /// It is the first parameter and the signature is what says so, since in the IR it is an
999    /// ordinary pointer and reads like one everywhere in the body. A function with a signature
1000    /// like that and no entry block has nothing to give back and no body to give it back from.
1001    fn sret(&self) -> Option<Value> {
1002        let first = self.source.signature().params.first()?;
1003        if !matches!(first.abi, Abi::Sret { .. }) {
1004            return None;
1005        }
1006        self.source[self.source.entry()?].params.first().copied()
1007    }
1008
1009    /// One `return` the convention has to write, as the place each value has to be in by the end.
1010    ///
1011    /// One pseudo per value, each a read constrained to a return register, which is what a return
1012    /// of one value already is and is the whole of what either does. The `ret` itself comes from
1013    /// the epilogue for both, long after this, because the frame has to be given back first.
1014    ///
1015    /// The two register files are counted separately, so a structure of a `double` and a `long`
1016    /// leaves the `double` in the first vector register and the `long` in the first integer one
1017    /// rather than in the second of either. That is the same walk `rucc_codegen::abi` makes on
1018    /// the other side of the call, which is what makes the two ends agree.
1019    ///
1020    /// A function whose answer went through memory gives back the address it was handed, in front
1021    /// of nothing else, because a signature that returns that way returns nothing else. That the
1022    /// caller already knows the address is not enough: it is allowed to read the register instead,
1023    /// and a caller that does gets whatever the allocator last left there. In a leaf function that
1024    /// is usually the right answer by accident, and one call in the body is enough to make it a
1025    /// wild pointer, which is why this is written rather than left to luck.
1026    ///
1027    /// Where everything goes is worked out before anything is written, so a return this cannot
1028    /// make leaves no half of one behind.
1029    /// Whether what a `return` gives back is the one value that goes back on the x87 stack.
1030    fn gives_back_x87(&self, inst: Inst) -> bool {
1031        let [value] = self.source[self.source[inst].args] else { return false };
1032        abi::on_the_stack(self.source[value].ty)
1033    }
1034
1035    fn returned(&mut self, inst: Inst) -> Result<(), Unsupported> {
1036        let values: Vec<Value> = self.source[self.source[inst].args].to_vec();
1037        let (mut ints, mut floats) = (0usize, 0usize);
1038        let mut parts = Vec::with_capacity(values.len() + 1);
1039        // An eighty bit value goes back on the x87 stack, which is where the convention says it is
1040        // and is the one place a value is left rather than put in a register. So the whole of the
1041        // return is an `fld` of its slot, and the stack it leaves the value on is not empty at the
1042        // `ret`, which is the one time in this file that is true and is what the convention asks
1043        // for. What comes after is the epilogue, which gives the frame back and touches nothing in
1044        // the unit.
1045        if let [value] = values[..] {
1046            let ty = self.source[value].ty;
1047            if abi::on_the_stack(ty) && self.sret().is_none() {
1048                let span = self.source.span(inst);
1049                let from = self.x87_slot(value);
1050                let from = self.through(from);
1051                self.x87_at("fld_t", span, from);
1052                return Ok(());
1053            }
1054        }
1055        for value in self.sret().into_iter().chain(values) {
1056            let ty = self.source[value].ty;
1057            let at = if crate::term::in_vector_file(ty) { &mut floats } else { &mut ints };
1058            // Why it cannot come back, and not only that it cannot. A type that travels nowhere
1059            // says so itself, and a type that travels perfectly well ran out of registers.
1060            let missing = abi::refuses(ty).unwrap_or(Missing::NoRoom);
1061            let name = abi::ret_of(ty, *at).ok_or(Unsupported::Returned { inst, missing })?;
1062            *at += 1;
1063            // The register is the target's answer and not one worked out here, the same as it is
1064            // for a return of one value, so that both halves of a pair and every rule that writes
1065            // half of one are reading the same table.
1066            let opcode = name.strip_prefix(PREFIX).expect("a machine instruction of this target");
1067            let form = x86_64::form(opcode).ok_or_else(|| self.unsupported(inst))?;
1068            let [desc] = form.operands() else { return Err(self.unsupported(inst)) };
1069            parts.push((self.names.intern(name), self.reg_of(value)?, *desc));
1070        }
1071
1072        let block = self.at.expect("a block is being filled");
1073        let span = self.source.span(inst);
1074        for (opcode, reg, desc) in parts {
1075            let operand = mir::Operand {
1076                reg,
1077                class: desc.class,
1078                role: desc.role,
1079                constraint: desc.constraint,
1080            };
1081            self.out.build(block, mir::Opcode::new(opcode)).at(span).operand(operand).finish();
1082        }
1083        Ok(())
1084    }
1085
1086    /// One `alloca`: the bytes it asks for go on the list the frame is laid out from, and the
1087    /// address of them is one instruction.
1088    ///
1089    /// The instruction is a `lea` off the stack pointer, which is the one register that reaches
1090    /// the frame in every function, and its displacement is left at nothing because there is no
1091    /// frame yet. Which instruction is waiting for which local is remembered, and
1092    /// [`crate::finish`] fills the numbers in after [`crate::frame::Frame`] has placed them.
1093    ///
1094    /// There is deliberately no rule for `alloca` and no name for one in [`crate::term`], and
1095    /// that is what stops it being folded into something else. An operand shown as the
1096    /// instruction that computed it is offered to the matcher by its name, so an `alloca` with no
1097    /// name is one no pattern can reach past, and the address it computes is always in a register
1098    /// by the time anything reads it.
1099    fn reserve(&mut self, inst: Inst) -> Result<(), Unsupported> {
1100        let data = &self.source[inst];
1101        // A variable length array carries the size it wants as an operand rather than in the
1102        // instruction, which is the whole of what tells the two apart here.
1103        if let Some(&size) = self.source[data.args].first() {
1104            return self.grow(inst, size);
1105        }
1106        let Extra::Mem(mem) = data.extra else { return Err(self.unsupported(inst)) };
1107        let info = self.source[mem];
1108        let size = u32::try_from(info.size)
1109            .map_err(|_| Unsupported::Dynamic { inst, growing: Growing::Huge })?;
1110        let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
1111
1112        // At least one, because the frame divides by the alignment and an object with no
1113        // alignment at all is one the front end had nothing to say about rather than one that may
1114        // go anywhere.
1115        let index = self.stack.locals.len();
1116        self.stack.locals.push(Local { size, align: info.align.max(1) });
1117
1118        let block = self.at.expect("a block is being filled");
1119        let reg = self.new_reg(result);
1120        let span = self.source.span(inst);
1121        let lea = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.lea)));
1122        let sp = mir::Operand::read(mir::Reg::physical(self.conv.stack_pointer), self.gpr);
1123        let made =
1124            self.out.build(block, lea).at(span).def(reg, self.gpr).mem(mir::Mem::at(sp)).finish();
1125        self.stack.addresses.push((made, index));
1126        Ok(())
1127    }
1128
1129    /// The other kind of `alloca`: one whose size the function does not know until it runs, which
1130    /// is what a variable length array is.
1131    ///
1132    /// Nothing about it is a slot the frame laid out, because the frame is laid out once and this
1133    /// happens as often as control reaches the declaration. The bytes come off the stack pointer
1134    /// where the declaration stands, which is two instructions:
1135    ///
1136    /// ```text
1137    ///   sub sp, bytes     the stack pointer moves down over the memory, which is what takes it
1138    ///   lea reg, [sp+n]   where the memory starts, which is above the outgoing argument area
1139    /// ```
1140    ///
1141    /// The displacement is left at nothing for the reason the constant kind leaves its own at
1142    /// nothing, and for a different number: that area belongs to the arguments of whatever this
1143    /// function calls, it stays at the bottom of the frame wherever the bottom has moved to, and
1144    /// how big it is is not known until every call in the function has been seen.
1145    ///
1146    /// The bytes are already a multiple of the stack pointer's alignment by the time they arrive,
1147    /// because [`crate::expand::rounds`] rounded them up in the IR, so nothing here has to mask the
1148    /// stack pointer afterwards and the stack pointer stays somewhere a call can be made from.
1149    ///
1150    /// Refused for an array wanting more alignment than the convention leaves the stack pointer
1151    /// with. Forcing that would be a second rounding of a register the frame already rounded, and
1152    /// after it no constant reaches the rest of the frame from anywhere. See `Growing` in
1153    /// [`crate::frame`].
1154    fn grow(&mut self, inst: Inst, size: Value) -> Result<(), Unsupported> {
1155        let data = &self.source[inst];
1156        let Extra::Mem(mem) = data.extra else { return Err(self.unsupported(inst)) };
1157        let info = self.source[mem];
1158        if info.align > self.conv.stack_align {
1159            return Err(Unsupported::Dynamic { inst, growing: Growing::Aligned });
1160        }
1161        let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
1162        let bytes = self.reg_of(size)?;
1163
1164        let block = self.at.expect("a block is being filled");
1165        let span = self.source.span(inst);
1166        let stack = mir::Reg::physical(self.conv.stack_pointer);
1167        let grow = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.grow)));
1168        self.out
1169            .build(block, grow)
1170            .at(span)
1171            .operand(mir::Operand::write(stack, self.gpr))
1172            .operand(mir::Operand::read(stack, self.gpr))
1173            .operand(mir::Operand::read(bytes, self.gpr))
1174            .finish();
1175
1176        let reg = self.new_reg(result);
1177        let lea = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.lea)));
1178        let sp = mir::Operand::read(stack, self.gpr);
1179        let made =
1180            self.out.build(block, lea).at(span).def(reg, self.gpr).mem(mir::Mem::at(sp)).finish();
1181        self.stack.dynamic.push(made);
1182        self.stack.grown_at.get_or_insert(inst);
1183        Ok(())
1184    }
1185
1186    /// Where the stack pointer is, kept so that something later can put it back.
1187    ///
1188    /// One move out of the stack pointer and one move into it, which is the whole of what the two
1189    /// halves are. What makes them worth writing is where the front end puts them: a scope holding
1190    /// a variable length array saves the stack pointer as it opens and puts it back as it closes,
1191    /// so a loop declaring one takes its bytes once round rather than once per iteration, and a
1192    /// jump out of the scope gives the bytes back on the way out.
1193    ///
1194    /// The value travels in an ordinary register the allocator hands out, so it may be spilled like
1195    /// any other, and a spill slot in a frame that grows is reached through the frame pointer,
1196    /// which is exactly the register that still means something after the stack pointer has moved.
1197    fn stack_pointer(&mut self, inst: Inst, into: bool) -> Result<(), Unsupported> {
1198        let data = &self.source[inst];
1199        let block = self.at.expect("a block is being filled");
1200        let span = self.source.span(inst);
1201        let stack = mir::Reg::physical(self.conv.stack_pointer);
1202        let mov = x86_64::FRAME.moves(self.gpr).expect("a class the target says how to move").mov;
1203        let mov = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{mov}")));
1204        let (write, read) = if into {
1205            let &saved = self.source[data.args].first().ok_or_else(|| self.unsupported(inst))?;
1206            (stack, self.reg_of(saved)?)
1207        } else {
1208            let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
1209            (self.new_reg(result), stack)
1210        };
1211        self.out
1212            .build(block, mov)
1213            .at(span)
1214            .operand(mir::Operand::write(write, self.gpr))
1215            .operand(mir::Operand::read(read, self.gpr))
1216            .finish();
1217        // Only the write is a move of the stack pointer, and it is the one that makes the frame a
1218        // growing one. A read of it in a function that never writes it back is a function that
1219        // asked where the stack was and did nothing with the answer.
1220        if into {
1221            self.stack.grown_at.get_or_insert(inst);
1222        }
1223        Ok(())
1224    }
1225
1226    /// Whether an instruction has an eighty bit float anywhere in it.
1227    ///
1228    /// Producing one and reading one are the same question here, because what makes one of these
1229    /// different from every other instruction is not the operation but where the value is. A
1230    /// `long double` is on the x87 stack while it is being worked on and in a frame slot the rest
1231    /// of the time, and neither of those is somewhere the operand of a rule could point.
1232    fn touches_x87(&self, inst: Inst) -> bool {
1233        let data = &self.source[inst];
1234        data.results().any(|value| on_x87(self.source[value].ty))
1235            || self.source[data.args].iter().any(|&arg| on_x87(self.source[arg].ty))
1236    }
1237
1238    /// Everything that happens to an eighty bit float, as the group of instructions it is.
1239    ///
1240    /// The first six move one, and every one of those is a load, a store, or a load and a store at
1241    /// two different formats, because that is the whole of what this machine converts with: the
1242    /// x87 has no instruction that turns one thing on its stack into another, so a widening is
1243    /// `fld` of the narrow format and a narrowing is `fstp` of it.
1244    ///
1245    /// The rest work on one, and they are here rather than in a rule for the same reason the six
1246    /// are. An add is a push, a push, the add and a pop, and what passes between those four is the
1247    /// top of a stack nothing allocates from, so there is no value in the middle of the group for
1248    /// a pattern to bind or a replacement to name. The comparison is the same shape with its last
1249    /// two instructions folded into one opcode, which is where the byte it produces comes from.
1250    ///
1251    /// Every group leaves the stack as empty as it found it, which is what `spec/10-backend.md`
1252    /// section 10.8 asks of one and is why nothing in this file has to track a depth: each push
1253    /// below is answered by a pop a line or two later, so no two groups can ever be looking at
1254    /// the same eight registers.
1255    fn x87(&mut self, inst: Inst) -> Result<(), Unsupported> {
1256        match self.source[inst].opcode {
1257            Opcode::Load => self.x87_load(inst),
1258            Opcode::Store => self.x87_store(inst),
1259            Opcode::FPExt => self.x87_widen(inst),
1260            Opcode::FPTrunc => self.x87_narrow(inst),
1261            Opcode::SIToFP => self.x87_from_signed(inst),
1262            Opcode::FPToSI => self.x87_to_signed(inst),
1263            Opcode::FAdd => self.x87_arith(inst, "fadd_p"),
1264            Opcode::FSub => self.x87_arith(inst, "fsubr_p"),
1265            Opcode::FMul => self.x87_arith(inst, "fmul_p"),
1266            Opcode::FDiv => self.x87_arith(inst, "fdivr_p"),
1267            Opcode::FNeg => self.x87_flip(inst),
1268            Opcode::FCmp => self.x87_compare(inst),
1269            Opcode::FConst => self.x87_const(inst),
1270            _ => Err(self.unsupported(inst)),
1271        }
1272    }
1273
1274    /// The eighty bit parameters of a block, copied out of the addresses an edge handed over and
1275    /// into slots of the block's own.
1276    ///
1277    /// What crosses an edge for a value of this type is an address, because the value is sixteen
1278    /// bytes of the frame and no register holds any of it. The block cannot keep that address: a
1279    /// second edge into the same block hands over a second one, and a read after the block would
1280    /// then be a read of whichever edge was taken rather than of one place. So the block has a
1281    /// slot per parameter and the bytes are copied into it here, which is the move on an edge that
1282    /// every other type gets from the allocator.
1283    ///
1284    /// Every load runs before every store and the stores run backwards, so all of the values are
1285    /// on the x87 stack at once and nothing reads a slot another one has already written. That
1286    /// costs nothing in the ordinary case of one parameter and is what makes the back edge of a
1287    /// loop that swaps two of these work. It is also the reason for the limit: the stack is eight
1288    /// deep, and a block with more of these than that is refused rather than copied in an order
1289    /// that could be wrong.
1290    fn settle(&mut self, block: Block, arriving: &[(Value, mir::Reg)]) -> Result<(), Unsupported> {
1291        let Some(&(first, _)) = arriving.first() else { return Ok(()) };
1292        if arriving.len() > X87_DEPTH {
1293            let ty = self.source[first].ty;
1294            return Err(Unsupported::Phi { block, count: arriving.len(), ty });
1295        }
1296        // A block parameter comes from no instruction, so what this points at is the first thing
1297        // in the block, which is where a reader looking for the copy would look.
1298        let first_inst = self.source.insts(block).next();
1299        let span = first_inst.map_or(Span::DUMMY, |it| self.source.span(it));
1300        for &(_, reg) in arriving {
1301            let from = self.through(reg);
1302            self.x87_at("fld_t", span, from);
1303        }
1304        for &(param, _) in arriving.iter().rev() {
1305            let into = self.x87_slot(param);
1306            let into = self.through(into);
1307            self.x87_at("fstp_t", span, into);
1308        }
1309        Ok(())
1310    }
1311
1312    /// The frame slot an eighty bit value lives in, as its address in a fresh register.
1313    ///
1314    /// The slot is the value's for the whole function and is taken the first time somebody asks.
1315    /// The address is worked out again every time, which is a `lea` per use and is deliberate: one
1316    /// address kept in a register from the definition to the last use would hold a general purpose
1317    /// register open across everything in between, and a function with a handful of these in it
1318    /// would spend its registers on addresses of things rather than on things.
1319    fn x87_slot(&mut self, value: Value) -> mir::Reg {
1320        // An argument of the function has a slot already and it is the caller's. The convention
1321        // puts the bytes in the argument area and hands over where they are, so the address that
1322        // arrived is the answer and no second copy of the value is made. Nothing ever writes to a
1323        // value of this type once it exists, so nothing writes to the caller's copy either. A
1324        // parameter of any other block is not this: what arrived there is an address a predecessor
1325        // chose, [`Lowering::settle`] has already copied the bytes out of it, and the slot those
1326        // bytes landed in is the one below.
1327        let entry = self.source.entry();
1328        if let (Def::Param { block, .. }, Some(reg)) =
1329            (self.source[value].def, self.regs[value.index()])
1330        {
1331            if entry == Some(block) {
1332                return reg;
1333            }
1334        }
1335        let index = match self.slots[value.index()] {
1336            Some(index) => index,
1337            None => {
1338                let index = self.stack.locals.len();
1339                self.stack.locals.push(Local { size: X87_BYTES, align: X87_BYTES });
1340                self.slots[value.index()] = Some(index);
1341                index
1342            }
1343        };
1344        let block = self.at.expect("a block is being filled");
1345        self.frame_address(block, index)
1346    }
1347
1348    /// The bytes a value crosses between a register and the x87 stack through, as their address
1349    /// in a fresh register.
1350    fn x87_crossing(&mut self) -> mir::Reg {
1351        let index = match self.crossing {
1352            Some(index) => index,
1353            None => {
1354                let index = self.stack.locals.len();
1355                self.stack.locals.push(Local { size: X87_CROSSING, align: X87_CROSSING });
1356                self.crossing = Some(index);
1357                index
1358            }
1359        };
1360        let block = self.at.expect("a block is being filled");
1361        self.frame_address(block, index)
1362    }
1363
1364    /// The two control words, as the address of the first of them in a fresh register.
1365    fn x87_control(&mut self) -> mir::Reg {
1366        let index = match self.control {
1367            Some(index) => index,
1368            None => {
1369                let index = self.stack.locals.len();
1370                self.stack.locals.push(Local { size: 4, align: 4 });
1371                self.control = Some(index);
1372                index
1373            }
1374        };
1375        let block = self.at.expect("a block is being filled");
1376        self.frame_address(block, index)
1377    }
1378
1379    /// An address held in a register, as the addressing mode that reaches it.
1380    fn through(&self, reg: mir::Reg) -> mir::Mem {
1381        mir::Mem::at(mir::Operand::read(reg, self.gpr))
1382    }
1383
1384    /// One instruction of a group, which names an address and nothing else.
1385    ///
1386    /// Every x87 instruction that moves a value is one of these. What it does to the stack is in
1387    /// the mnemonic rather than in an operand, so there is no register to write down and no
1388    /// register the allocator gets a say in.
1389    fn x87_at(&mut self, name: &str, span: Span, at: mir::Mem) {
1390        let block = self.at.expect("a block is being filled");
1391        let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
1392        self.out.build(block, opcode).at(span).mem(at).finish();
1393    }
1394
1395    /// One instruction of a group that names nothing at all.
1396    ///
1397    /// The arithmetic is these. Both of an add's operands are already on the stack when it runs
1398    /// and so is where the answer goes, and the stack is not somewhere an instruction says, so
1399    /// `faddp` has an argument in the assembler's syntax and nothing here for the argument to come
1400    /// from. What it works on is which two pushes came before it, which is a fact about the order
1401    /// of the group and is why the group is written in one place.
1402    fn x87_only(&mut self, name: &str, span: Span) {
1403        let block = self.at.expect("a block is being filled");
1404        let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
1405        self.out.build(block, opcode).at(span).finish();
1406    }
1407
1408    /// A `load` of a `long double`: onto the stack from where it was, and off it into the slot.
1409    ///
1410    /// Two instructions rather than the two general purpose moves the same sixteen bytes would
1411    /// take, because `fld` and `fstp` at this format neither convert nor look: the value goes on
1412    /// in the format it was already in and comes back off in it, so a signalling NaN stays one
1413    /// and nothing is raised. Which is what makes this a copy at all.
1414    fn x87_load(&mut self, inst: Inst) -> Result<(), Unsupported> {
1415        let (args, result) = self.ends(inst)?;
1416        let &address = args.first().ok_or_else(|| self.unsupported(inst))?;
1417        let span = self.source.span(inst);
1418        let from = self.reg_of(address)?;
1419        let from = self.through(from);
1420        let into = self.x87_slot(result);
1421        let into = self.through(into);
1422        self.x87_at("fld_t", span, from);
1423        self.x87_at("fstp_t", span, into);
1424        Ok(())
1425    }
1426
1427    /// A `store` of a `long double`: the same pair the other way round.
1428    fn x87_store(&mut self, inst: Inst) -> Result<(), Unsupported> {
1429        let args = self.source[self.source[inst].args].to_vec();
1430        let [value, address] = args[..] else { return Err(self.unsupported(inst)) };
1431        let span = self.source.span(inst);
1432        let from = self.x87_slot(value);
1433        let from = self.through(from);
1434        let into = self.reg_of(address)?;
1435        let into = self.through(into);
1436        self.x87_at("fld_t", span, from);
1437        self.x87_at("fstp_t", span, into);
1438        Ok(())
1439    }
1440
1441    /// A `float`, a `double` or an integer becoming a `long double`.
1442    ///
1443    /// Through memory, because the x87 reads memory and nothing else: the value is in a register
1444    /// the machine has and the unit has no way to be handed one, so it is written to the crossing
1445    /// bytes and loaded back at the format that widens it. Every one of these is exact. Sixty four
1446    /// bits of significand and fifteen of exponent hold every `float`, every `double` and every
1447    /// sixty four bit integer outright, so none of the four can round and none can raise.
1448    fn x87_across(
1449        &mut self,
1450        inst: Inst,
1451        put: &'static str,
1452        class: RegClass,
1453        get: &'static str,
1454    ) -> Result<(), Unsupported> {
1455        let (args, result) = self.ends(inst)?;
1456        let &source = args.first().ok_or_else(|| self.unsupported(inst))?;
1457        let span = self.source.span(inst);
1458        let value = self.reg_of(source)?;
1459        let across = self.x87_crossing();
1460        let across = self.through(across);
1461        let into = self.x87_slot(result);
1462        let into = self.through(into);
1463
1464        let block = self.at.expect("a block is being filled");
1465        let store = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{put}")));
1466        self.out.build(block, store).at(span).uses(value, class).mem(across).finish();
1467        self.x87_at(get, span, across);
1468        self.x87_at("fstp_t", span, into);
1469        Ok(())
1470    }
1471
1472    /// A `long double` becoming a `float`, a `double` or an integer.
1473    ///
1474    /// Through memory for the reason above and in the same three instructions backwards. The two
1475    /// that go to a float round to nearest, which is what the control word says unless somebody
1476    /// has changed it and is what C wants. The two that go to an integer do not, which is why they
1477    /// do not come here.
1478    fn x87_back(
1479        &mut self,
1480        inst: Inst,
1481        put: &'static str,
1482        get: &'static str,
1483        class: RegClass,
1484    ) -> Result<(), Unsupported> {
1485        let (args, result) = self.ends(inst)?;
1486        let &source = args.first().ok_or_else(|| self.unsupported(inst))?;
1487        let span = self.source.span(inst);
1488        let from = self.x87_slot(source);
1489        let from = self.through(from);
1490        let across = self.x87_crossing();
1491        let across = self.through(across);
1492
1493        self.x87_at("fld_t", span, from);
1494        self.x87_at(put, span, across);
1495        let block = self.at.expect("a block is being filled");
1496        let reg = self.new_reg(result);
1497        let load = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{get}")));
1498        self.out.build(block, load).at(span).def(reg, class).mem(across).finish();
1499        Ok(())
1500    }
1501
1502    /// An `fpext` up to a `long double`, which is the only direction this machine has one in.
1503    fn x87_widen(&mut self, inst: Inst) -> Result<(), Unsupported> {
1504        let sse = self.conv.sse_class;
1505        match self.source[self.narrow(inst)?].ty.bits() {
1506            32 => self.x87_across(inst, "movss_mr", sse, "fld_s"),
1507            64 => self.x87_across(inst, "movsd_mr", sse, "fld_l"),
1508            _ => Err(self.unsupported(inst)),
1509        }
1510    }
1511
1512    /// An `fptrunc` down from a `long double`, which is the other direction of the same.
1513    fn x87_narrow(&mut self, inst: Inst) -> Result<(), Unsupported> {
1514        let sse = self.conv.sse_class;
1515        let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
1516        match self.source[result].ty.bits() {
1517            32 => self.x87_back(inst, "fstp_s", "movss_rm", sse),
1518            64 => self.x87_back(inst, "fstp_l", "movsd_rm", sse),
1519            _ => Err(self.unsupported(inst)),
1520        }
1521    }
1522
1523    /// A `sitofp` up to a `long double`.
1524    ///
1525    /// Thirty two bits and sixty four, and nothing narrower, because C widens an integer to `int`
1526    /// before it converts one and the front end writes that widening down. An unsigned integer is
1527    /// not here at all: `fild` reads its operand as signed, so a value above the signed range
1528    /// comes back short by two to the sixty fourth and has to be added back, which is arithmetic
1529    /// rather than a move and waits with the rest of it.
1530    fn x87_from_signed(&mut self, inst: Inst) -> Result<(), Unsupported> {
1531        let gpr = self.gpr;
1532        match self.source[self.narrow(inst)?].ty.bits() {
1533            32 => self.x87_across(inst, "mov_mr_32", gpr, "fild_l"),
1534            64 => self.x87_across(inst, "mov_mr_64", gpr, "fild_ll"),
1535            _ => Err(self.unsupported(inst)),
1536        }
1537    }
1538
1539    /// An `fptosi` down from a `long double`, which is the one conversion here with no single
1540    /// instruction behind it.
1541    ///
1542    /// C cuts towards zero and the unit rounds the way its control word says, so the store that
1543    /// takes the value off the stack is wrapped in the control word being saved, changed and put
1544    /// back. Five instructions around the one that does the work, and three more moving the word
1545    /// through a register, because this machine has no way to OR a constant into memory at this
1546    /// width. The unit has a shorter answer in `fisttp`, and `spec/10-backend.md` section 10.8
1547    /// says why it is not used: it is SSE3, the x86-64 baseline is not, and there is nothing here
1548    /// that can gate an instruction on a feature yet.
1549    fn x87_to_signed(&mut self, inst: Inst) -> Result<(), Unsupported> {
1550        let (args, result) = self.ends(inst)?;
1551        let &source = args.first().ok_or_else(|| self.unsupported(inst))?;
1552        let (put, get) = match self.source[result].ty.bits() {
1553            32 => ("fistp_l", "mov_rm_32"),
1554            64 => ("fistp_ll", "mov_rm_64"),
1555            _ => return Err(self.unsupported(inst)),
1556        };
1557        let span = self.source.span(inst);
1558        let gpr = self.gpr;
1559        let from = self.x87_slot(source);
1560        let from = self.through(from);
1561        let across = self.x87_crossing();
1562        let across = self.through(across);
1563        let control = self.x87_control();
1564        let saved = self.through(control).plus(0);
1565        let cut = self.through(control).plus(2);
1566
1567        // The word the unit has now, into the first of the two slots and into a register, with the
1568        // rounding field turned to truncate on the way to the second.
1569        self.x87_at("fnstcw", span, saved);
1570        let block = self.at.expect("a block is being filled");
1571        let was = self.out.new_vreg(gpr);
1572        let read = mir::Opcode::new(self.names.intern("x64.mov_rm_16"));
1573        self.out.build(block, read).at(span).def(was, gpr).mem(saved).finish();
1574        let now = self.out.new_vreg(gpr);
1575        let set = mir::Opcode::new(self.names.intern("x64.or_ri_16"));
1576        // Two address, which is written out here rather than taken from the two shorthands
1577        // because the shorthands leave an operand unconstrained: this machine ORs into the
1578        // register it read, so the two have to be the same one and only the constraint says so.
1579        self.out
1580            .build(block, set)
1581            .at(span)
1582            .operand(mir::Operand::write(now, gpr).with(Constraint::Reuse(1)))
1583            .operand(mir::Operand::read(was, gpr))
1584            .imm(X87_TRUNCATE)
1585            .finish();
1586        let write = mir::Opcode::new(self.names.intern("x64.mov_mr_16"));
1587        self.out.build(block, write).at(span).uses(now, gpr).mem(cut).finish();
1588
1589        // The conversion itself, under the changed word, and then the word the unit had put back
1590        // before anything else runs.
1591        self.x87_at("fldcw", span, cut);
1592        self.x87_at("fld_t", span, from);
1593        self.x87_at(put, span, across);
1594        self.x87_at("fldcw", span, saved);
1595
1596        let block = self.at.expect("a block is being filled");
1597        let reg = self.new_reg(result);
1598        let load = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{get}")));
1599        self.out.build(block, load).at(span).def(reg, gpr).mem(across).finish();
1600        Ok(())
1601    }
1602
1603    /// A constant of this type, as the bits of it written into its slot.
1604    ///
1605    /// No x87 instruction at all, which is the surprise here. A slot holding an eighty bit value is
1606    /// the value, so a constant is ten bytes put where the value lives, and the unit never has to
1607    /// see it: whatever reads it will `fld` it out of the slot the way it reads any other one.
1608    ///
1609    /// Ten bytes in two goes, because the machine stores eight at a time and there is no store of
1610    /// an immediate to memory, so each half is put in a register first. The six bytes above the ten
1611    /// are left alone, since nothing reads them: they are the padding that makes the type sixteen
1612    /// wide and they are unspecified in the psABI rather than zero.
1613    ///
1614    /// The other way is a constant pool, an `fldt` of a symbol, and a relocation, which is what a
1615    /// compiler with somewhere to put a literal does. This back end has nowhere to put one yet, and
1616    /// four instructions in the frame is what that costs until it does.
1617    fn x87_const(&mut self, inst: Inst) -> Result<(), Unsupported> {
1618        let Extra::Imm(imm) = self.source[inst].extra else { return Err(self.unsupported(inst)) };
1619        let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
1620        let bits = self.source[imm].bits();
1621        let span = self.source.span(inst);
1622        let gpr = self.gpr;
1623        let slot = self.x87_slot(result);
1624        let low = self.through(slot).plus(0);
1625        let high = self.through(slot).plus(8);
1626
1627        let block = self.at.expect("a block is being filled");
1628        for (bytes, at, into) in
1629            [(bits as u64 as i64, low, "64"), (((bits >> 64) & 0xffff) as i64, high, "16")]
1630        {
1631            let held = self.out.new_vreg(gpr);
1632            let put = mir::Opcode::new(self.names.intern(&format!("{PREFIX}mov_ri_{into}")));
1633            self.out.build(block, put).at(span).def(held, gpr).imm(bytes).finish();
1634            let store = mir::Opcode::new(self.names.intern(&format!("{PREFIX}mov_mr_{into}")));
1635            self.out.build(block, store).at(span).uses(held, gpr).mem(at).finish();
1636        }
1637        Ok(())
1638    }
1639
1640    /// One arithmetic instruction on two eighty bit values, as the four it takes.
1641    ///
1642    /// The left operand is pushed first and the right one on top of it, so the left ends up
1643    /// underneath and the answer wanted is the one below against the top in that order. Which of
1644    /// the two mnemonics computes that is a question about the spelling rather than about the
1645    /// machine, and the two spellings disagree. Intel's `FSUBP ST(i), ST(0)` is `ST(i) - ST(0)`
1646    /// and is `DE E8+i`, and AT&T's `fsubp` is `DE E0+i`, which is the other subtraction. This
1647    /// compiler writes AT&T and encodes what gas encodes, so what it asks for here is `fsubr_p`
1648    /// and `fdivr_p`, and the `r` is not a reversal of anything the code generator decided.
1649    ///
1650    /// An addition and a multiplication have one form each and do not care, which is why a test
1651    /// that reads the mnemonic back would not have caught this and one that computes a subtraction
1652    /// and checks the answer does.
1653    ///
1654    /// The answer is left where the deeper of the two was and the shallower is gone, which is what
1655    /// the `p` on the mnemonic means, so one push has already been paid back by the time the
1656    /// `fstp` runs and the stack is level again after it.
1657    ///
1658    /// Nothing here is folded and nothing is reused. Two values that are the same value get two
1659    /// pushes of the same slot, and an operand that was just computed is read back out of the slot
1660    /// it was written to rather than left on the stack, which costs a store and a load per
1661    /// instruction in an expression. Keeping a partial result on the stack across the next
1662    /// instruction's operands means knowing how deep the stack is at every point in the block, and
1663    /// that is a different thing from writing a group.
1664    fn x87_arith(&mut self, inst: Inst, with: &'static str) -> Result<(), Unsupported> {
1665        let (args, result) = self.ends(inst)?;
1666        let [left, right] = args[..] else { return Err(self.unsupported(inst)) };
1667        let span = self.source.span(inst);
1668        let left = self.x87_slot(left);
1669        let left = self.through(left);
1670        let right = self.x87_slot(right);
1671        let right = self.through(right);
1672        let into = self.x87_slot(result);
1673        let into = self.through(into);
1674        self.x87_at("fld_t", span, left);
1675        self.x87_at("fld_t", span, right);
1676        self.x87_only(with, span);
1677        self.x87_at("fstp_t", span, into);
1678        Ok(())
1679    }
1680
1681    /// A negation, which is a push, the sign bit turned over and a pop.
1682    ///
1683    /// `fchs` does not read the value as a number, so this is right for a zero, for an infinity
1684    /// and for a NaN, and it raises nothing on any of them. Which is what C asks of a negation and
1685    /// is not what subtracting from zero would give: `0.0L - x` is a different answer at a
1686    /// negative zero and a signalling one at a NaN.
1687    fn x87_flip(&mut self, inst: Inst) -> Result<(), Unsupported> {
1688        let (args, result) = self.ends(inst)?;
1689        let &source = args.first().ok_or_else(|| self.unsupported(inst))?;
1690        let span = self.source.span(inst);
1691        let from = self.x87_slot(source);
1692        let from = self.through(from);
1693        let into = self.x87_slot(result);
1694        let into = self.through(into);
1695        self.x87_at("fld_t", span, from);
1696        self.x87_only("fchs", span);
1697        self.x87_at("fstp_t", span, into);
1698        Ok(())
1699    }
1700
1701    /// A comparison of two eighty bit values, as the two pushes and the one opcode that reads them.
1702    ///
1703    /// The right operand is pushed first and the left one on top of it, which is the other way
1704    /// round from the arithmetic and is because `fucomip` asks about the top against what is under
1705    /// it: the comparison this machine can do is the top's, so the value the predicate is about
1706    /// has to be the top. The pop that gets the loser off the stack and the byte that reads the
1707    /// flags are both inside the opcode, since what passes between those and the comparison is the
1708    /// flags and the flags are not something anything here can name.
1709    ///
1710    /// Which of the ten opcodes, and which way round, is the same table the vector comparisons
1711    /// match against in `rules/x86-64.rules`, and it has to stay the same table: a predicate that
1712    /// picked a different condition here than there would be a `long double` comparison that
1713    /// disagreed with the `double` comparison of the same two numbers, which is the one thing a
1714    /// wider format is not allowed to do.
1715    ///
1716    /// The always false and the always true are refused rather than folded into a constant,
1717    /// because a comparison this machine never has to do is one the optimizer should have removed
1718    /// and an instruction here that quietly agreed with it would hide that it did not.
1719    fn x87_compare(&mut self, inst: Inst) -> Result<(), Unsupported> {
1720        let Extra::FloatPred(pred) = self.source[inst].extra else {
1721            return Err(self.unsupported(inst));
1722        };
1723        let (args, result) = self.ends(inst)?;
1724        let [left, right] = args[..] else { return Err(self.unsupported(inst)) };
1725        // Two of the fourteen need a second byte and an instruction to put the two together,
1726        // because they are two conditions at once: an ordered equal is equal and not unordered,
1727        // and an unordered not equal is either. The opcode carries all of that and says here only
1728        // that it writes somewhere else as well.
1729        let (name, reversed, both) = match pred {
1730            FloatPred::Ogt => ("fucomip_set_a", false, false),
1731            FloatPred::Oge => ("fucomip_set_ae", false, false),
1732            FloatPred::Olt => ("fucomip_set_a", true, false),
1733            FloatPred::Ole => ("fucomip_set_ae", true, false),
1734            FloatPred::One => ("fucomip_set_ne", false, false),
1735            FloatPred::Ord => ("fucomip_set_np", false, false),
1736            FloatPred::Uno => ("fucomip_set_p", false, false),
1737            FloatPred::Ueq => ("fucomip_set_e", false, false),
1738            FloatPred::Ult => ("fucomip_set_b", false, false),
1739            FloatPred::Ule => ("fucomip_set_be", false, false),
1740            FloatPred::Ugt => ("fucomip_set_b", true, false),
1741            FloatPred::Uge => ("fucomip_set_be", true, false),
1742            FloatPred::Oeq => ("fucomip_set_e_and_np", false, true),
1743            FloatPred::Une => ("fucomip_set_ne_or_p", false, true),
1744            FloatPred::False | FloatPred::True => return Err(self.unsupported(inst)),
1745        };
1746        let (top, under) = if reversed { (right, left) } else { (left, right) };
1747
1748        let span = self.source.span(inst);
1749        let gpr = self.gpr;
1750        let under = self.x87_slot(under);
1751        let under = self.through(under);
1752        let top = self.x87_slot(top);
1753        let top = self.through(top);
1754        self.x87_at("fld_t", span, under);
1755        self.x87_at("fld_t", span, top);
1756
1757        let block = self.at.expect("a block is being filled");
1758        let reg = self.new_reg(result);
1759        // Taken before the instruction is started rather than inside it, since both come from the
1760        // same function being built and only one thing at a time may be adding to it.
1761        let spare = both.then(|| self.out.new_vreg(gpr));
1762        let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
1763        let mut build = self.out.build(block, opcode).at(span).def(reg, gpr);
1764        if let Some(spare) = spare {
1765            build = build.def(spare, gpr);
1766        }
1767        build.finish();
1768        Ok(())
1769    }
1770
1771    /// The operands and the one result of an instruction that has exactly one.
1772    fn ends(&self, inst: Inst) -> Result<(&'a [Value], Value), Unsupported> {
1773        let data = &self.source[inst];
1774        let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
1775        Ok((&self.source[data.args], result))
1776    }
1777
1778    /// The operand of a conversion, which is the end of it that is not the `long double`.
1779    fn narrow(&self, inst: Inst) -> Result<Value, Unsupported> {
1780        let args = &self.source[self.source[inst].args];
1781        args.first().copied().ok_or_else(|| self.unsupported(inst))
1782    }
1783
1784    /// One `va_start`, as the four fields of the list it was handed.
1785    ///
1786    /// Two of them are numbers this already knows, and each costs an instruction to put in a
1787    /// register before it can be stored, because the machine here has no store of an immediate to
1788    /// memory. The other two are addresses in the frame, and each is a `lea` [`crate::finish`]
1789    /// finishes: the save area is one of the function's own stack objects, and the caller's
1790    /// argument area is where the parameters that had no register came from, which is the same
1791    /// place and the same fixup a parameter past the sixth already uses.
1792    ///
1793    /// What is written is exactly the four fields [`crate::varargs`] describes, in the order they
1794    /// are laid out, so that reading this beside that table is the whole of the check.
1795    fn va_start(&mut self, inst: Inst) -> Result<(), Unsupported> {
1796        let Some(&list) = self.source[self.source[inst].args].first() else {
1797            return Err(self.unsupported(inst));
1798        };
1799        let started = self.varargs.ok_or_else(|| self.unsupported(inst))?;
1800        let list = self.reg_of(list)?;
1801        let block = self.at.expect("a block is being filled");
1802        let span = self.source.span(inst);
1803
1804        for (at, count) in
1805            [(varargs::GP_OFFSET, started.integers), (varargs::FP_OFFSET, started.floats)]
1806        {
1807            let held = self.out.new_vreg(self.gpr);
1808            let load = mir::Opcode::new(self.names.intern("x64.mov_ri_32"));
1809            self.out.build(block, load).at(span).def(held, self.gpr).imm(i64::from(count)).finish();
1810
1811            let store = mir::Opcode::new(self.names.intern("x64.mov_mr_32"));
1812            let mem = self.field(list, at);
1813            self.out.build(block, store).at(span).uses(held, self.gpr).mem(mem).finish();
1814        }
1815
1816        // The first argument the signature did not name, which is as far up the caller's argument
1817        // area as the ones it did name reached. Nothing here knows where that area is, so the
1818        // distance is recorded the way a parameter read out of it is and finished with it.
1819        let overflow = self.out.new_vreg(self.gpr);
1820        let lea = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.lea)));
1821        let sp = mir::Operand::read(mir::Reg::physical(self.conv.stack_pointer), self.gpr);
1822        let made = self
1823            .out
1824            .build(block, lea)
1825            .at(span)
1826            .def(overflow, self.gpr)
1827            .mem(mir::Mem::at(sp))
1828            .finish();
1829        self.stack.arguments.push((made, started.incoming));
1830
1831        let save = self.frame_address(block, started.save);
1832        for (at, held) in [(varargs::OVERFLOW, overflow), (varargs::SAVE_AREA, save)] {
1833            let store = mir::Opcode::new(self.names.intern("x64.mov_mr_64"));
1834            let mem = self.field(list, at);
1835            self.out.build(block, store).at(span).uses(held, self.gpr).mem(mem).finish();
1836        }
1837        Ok(())
1838    }
1839
1840    /// One field of a list, as the addressing mode that reaches it.
1841    fn field(&self, list: mir::Reg, at: i64) -> mir::Mem {
1842        let base = mir::Operand::read(list, self.gpr);
1843        mir::Mem::at(base).plus(i32::try_from(at).expect("a field of a list is a small offset"))
1844    }
1845
1846    /// The address of a name: one `lea` off the instruction pointer, with the name on it.
1847    ///
1848    /// The same instruction an `alloca` gets and for a related reason. An address that is not in
1849    /// the program is a `lea` of an addressing mode that names no register, and the mode carries
1850    /// the symbol so that [`rucc_asm`] can write it relative to `%rip` and leave the relocation
1851    /// for the assembler. Both halves of that already existed: the printer writes `sym(%rip)` and
1852    /// the encoder emits the relocation, because a call to a name the file does not define needed
1853    /// them first.
1854    ///
1855    /// One `mov` and not one `lea` when the name is one [`Elsewhere`] holds, because the distance
1856    /// the `lea` adds to the instruction pointer is a number only a link that puts the name in
1857    /// this program can work out, and the address of a function this file merely declares is not
1858    /// such a number. The load reads the address out of the slot the linker fills in instead. The
1859    /// linker turns it back into the `lea` when the name turns out to have been here all along,
1860    /// so this is not slower in the case that was already right.
1861    ///
1862    /// There is deliberately no name for this in [`crate::term`], which is what stops the address
1863    /// being folded into the instruction that reads it. Folding it is the right thing to do and
1864    /// is what turns a load of a global from two instructions into one, but it is a separate
1865    /// question about addressing modes and issue #282 is it. Until then the address is in a
1866    /// register before anything uses it, which is correct and one instruction longer.
1867    ///
1868    /// What this does not do is give the name anything to refer to. A module carries its globals
1869    /// and nothing writes them out, so a file that defines the variable it reads compiles to a
1870    /// reference the linker cannot resolve. Issue #293 is the other half.
1871    ///
1872    /// A thread-local variable is neither of the two above and is [`Self::thread_address`].
1873    fn address_of(&mut self, inst: Inst) -> Result<(), Unsupported> {
1874        let data = &self.source[inst];
1875        let Extra::Symbol(symbol) = data.extra else { return Err(self.unsupported(inst)) };
1876        let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
1877        if self.elsewhere.thread(symbol) {
1878            return self.thread_address(inst, symbol, result);
1879        }
1880
1881        let block = self.at.expect("a block is being filled");
1882        let reg = self.new_reg(result);
1883        let span = self.source.span(inst);
1884        let (mnemonic, mem) = if self.elsewhere.holds(symbol) {
1885            (GOT_LOAD, mir::Mem::got(symbol))
1886        } else {
1887            (x86_64::FRAME.lea, mir::Mem::of(symbol))
1888        };
1889        let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{mnemonic}")));
1890        self.out.build(block, opcode).at(span).def(reg, self.gpr).mem(mem).finish();
1891        Ok(())
1892    }
1893
1894    /// The address of a thread-local variable, which is this thread's copy of it.
1895    ///
1896    /// Neither instruction the ordinary case writes would mean anything here. There is no distance
1897    /// to the variable for a `lea` to add, because there is no variable: there is one copy of it per
1898    /// thread and they are at different addresses, so a link asked for the distance to the name
1899    /// refuses rather than picking one. And there is no address for a table slot to hold either, for
1900    /// the same reason.
1901    ///
1902    /// What is the same in every thread is where the variable sits inside the block of storage a
1903    /// thread gets, so that offset is what the link writes down, and the address of the running
1904    /// thread's block is what turns it into an address. x86-64 keeps that address in `%fs`, at the
1905    /// front of the block, so the whole of this is three instructions:
1906    ///
1907    /// ```text
1908    /// movq  x@gottpoff(%rip), %off   # how far into the block x sits, which the link fills in
1909    /// movq  %fs:0, %tp               # where this thread's block is, which only the machine knows
1910    /// addq  %tp, %off                # this thread's copy of x
1911    /// ```
1912    ///
1913    /// That is the initial exec model. It is one instruction longer than what gcc writes at `-O2`
1914    /// in an executable, which folds the addition into the instruction that uses the address, and
1915    /// the difference is issue #282 rather than anything about threads: nothing here folds an
1916    /// address into its reader yet. The link relaxes the first instruction into an immediate when it
1917    /// is making an executable, since it lays the blocks out and therefore knows the number, so the
1918    /// table slot costs nothing in the case that is common.
1919    ///
1920    /// It is not the most general model. A library loaded by `dlopen` gets its storage after the
1921    /// program is already running, and the block this reaches was laid out before it started, so
1922    /// the loader has to find room in that block for the library's variables. glibc keeps a little
1923    /// spare room for exactly this and a library that fits in it loads and runs; one that does not
1924    /// fails to load, with a message saying so. The model with no such limit calls `__tls_get_addr`
1925    /// and is what gcc writes under `-fPIC` by default, and it is issue #1104.
1926    ///
1927    /// So this is the model gcc writes under `-ftls-model=initial-exec`: right for an executable,
1928    /// right for a library the program is linked against, and a load that either works or is
1929    /// refused out loud for a library something opens later. What it is never is quietly wrong.
1930    fn thread_address(
1931        &mut self,
1932        inst: Inst,
1933        symbol: Symbol,
1934        result: Value,
1935    ) -> Result<(), Unsupported> {
1936        let block = self.at.expect("a block is being filled");
1937        let span = self.source.span(inst);
1938        let gpr = self.gpr;
1939        let load = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{GOT_LOAD}")));
1940
1941        let offset = self.out.new_vreg(gpr);
1942        self.out
1943            .build(block, load)
1944            .at(span)
1945            .def(offset, gpr)
1946            .mem(mir::Mem::thread(symbol))
1947            .finish();
1948        // The front of the block, which is the one thing on this machine that no instruction can
1949        // work out: `%fs` is not a register a program can read, and what it points at is a word
1950        // holding its own address, so reading through it at zero is how the address is come by.
1951        let pointer = self.out.new_vreg(gpr);
1952        let at = mir::Mem::in_segment(Segment::Fs, 0);
1953        self.out.build(block, load).at(span).def(pointer, gpr).mem(at).finish();
1954
1955        // Two address, spelled out for the reason `x87_to_int` gives: this machine adds into the
1956        // register it read, and only the constraint says the two are the same one.
1957        let reg = self.new_reg(result);
1958        let add = mir::Opcode::new(self.names.intern(&format!("{PREFIX}add_rr_64")));
1959        self.out
1960            .build(block, add)
1961            .at(span)
1962            .operand(mir::Operand::write(reg, gpr).with(Constraint::Reuse(1)))
1963            .operand(mir::Operand::read(offset, gpr))
1964            .operand(mir::Operand::read(pointer, gpr))
1965            .finish();
1966        Ok(())
1967    }
1968
1969    /// `&&label`, GNU's address of a label, which is the same `lea` a global gets against a place
1970    /// in this same function.
1971    ///
1972    /// What the two have in common is the whole of the instruction: an address worked out from
1973    /// where the instruction is, which is what `(%rip)` means and is the only way this compiler
1974    /// reaches anything. What they do not have in common is what fills the four bytes in. A
1975    /// global is a name, so the number is a relocation and the linker writes it. A block is a
1976    /// place in this function, so both ends are in one section and the number is known as soon as
1977    /// the blocks have been laid out, which is why `rucc_asm` fills it in the way it fills in a
1978    /// jump rather than leaving a relocation behind.
1979    ///
1980    /// Nothing here says the block is one control can arrive at. That is said by the
1981    /// [`Opcode::IndirectBr`] that reads the address, which lists every block it can arrive at,
1982    /// and by nothing else: an address on its own is a number.
1983    fn block_address(&mut self, inst: Inst) -> Result<(), Unsupported> {
1984        let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
1985        let Some(call) = self.source.successors(inst).next() else {
1986            return Err(self.unsupported(inst));
1987        };
1988        let block = self.at.expect("a block is being filled");
1989        let reg = self.new_reg(result);
1990        let span = self.source.span(inst);
1991        let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.lea)));
1992        let mem = mir::Mem::block(self.out_block(call.block));
1993        self.out.build(block, opcode).at(span).def(reg, self.gpr).mem(mem).finish();
1994        Ok(())
1995    }
1996
1997    /// `goto *p`, GNU's computed goto, which is a jump through a register.
1998    ///
1999    /// Where it goes is not written here and cannot be. Every block it can arrive at is on the
2000    /// block this ends, the way every other arm is, and which of them the address holds is decided
2001    /// while the program runs. So this is one instruction with one operand, and the arms are
2002    /// copied across by [`Self::edges`] like anybody else's.
2003    fn indirect_branch(&mut self, inst: Inst) -> Result<(), Unsupported> {
2004        let data = &self.source[inst];
2005        let &address = self.source[data.args].first().ok_or_else(|| self.unsupported(inst))?;
2006        let reg = self.reg_of(address)?;
2007        let block = self.at.expect("a block is being filled");
2008        let span = self.source.span(inst);
2009        let name = x86_64::BRANCH.indirect;
2010        let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
2011        self.out.build(block, opcode).at(span).operand(mir::Operand::read(reg, self.gpr)).finish();
2012        Ok(())
2013    }
2014
2015    /// `__builtin_thread_pointer`, which is the front of the block [`Self::thread_address`] adds
2016    /// an offset to.
2017    ///
2018    /// The same one instruction, on its own this time and with nothing to add to it. A program
2019    /// writes this when what it wants is a number that is different in every thread and cheap to
2020    /// come by, rather than a variable of its own in the block, so there is no relocation here and
2021    /// no name for the link to resolve.
2022    fn thread_pointer(&mut self, inst: Inst) -> Result<(), Unsupported> {
2023        let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
2024        let block = self.at.expect("a block is being filled");
2025        let span = self.source.span(inst);
2026        let reg = self.new_reg(result);
2027        let load = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{GOT_LOAD}")));
2028        let at = mir::Mem::in_segment(Segment::Fs, 0);
2029        self.out.build(block, load).at(span).def(reg, self.gpr).mem(at).finish();
2030        Ok(())
2031    }
2032
2033    /// A conversion that converts nothing: the result is the operand under another type.
2034    ///
2035    /// `ptrtoint` and `inttoptr` at one width are the whole of this. An address on this machine is
2036    /// an integer as wide as the machine addresses, so a cast between the two changes what the
2037    /// type system calls the value and changes nothing about the value, and the register holding
2038    /// it is the register that already held it. The front end never writes either of them at any
2039    /// other width, because it widens or narrows around the cast rather than through it, so the
2040    /// two widths disagreeing here means the IR came from somewhere else and is refused rather
2041    /// than guessed at.
2042    ///
2043    /// Reading the operand first is what materializes it when it is a constant, which is the case
2044    /// that matters: a null pointer is an `inttoptr` of zero, and that zero has to reach a
2045    /// register before anything can call it an address.
2046    fn rename(&mut self, inst: Inst) -> Result<(), Unsupported> {
2047        let data = &self.source[inst];
2048        let [arg] = self.source[data.args] else { return Err(self.unsupported(inst)) };
2049        let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
2050        if !self.is_address_width(self.source[arg].ty)
2051            || !self.is_address_width(self.source[result].ty)
2052        {
2053            return Err(self.unsupported(inst));
2054        }
2055        let reg = self.reg_of(arg)?;
2056        self.regs[result.index()] = Some(reg);
2057        Ok(())
2058    }
2059
2060    /// One barrier, which on this machine is one instruction at the strongest ordering and no
2061    /// instruction at all at every other one.
2062    ///
2063    /// x86-64 is total store order, so the only reordering the machine does is a store followed by
2064    /// a load of a different address, and the only ordering that forbids that is sequential
2065    /// consistency. An acquire, a release and an acquire release fence are therefore already true
2066    /// of every program running here, and what a program wanted from writing one is that the
2067    /// compiler not move memory accesses across it. The optimizer has finished by the time this
2068    /// runs and nothing below reorders one access past another, so the constraint is already
2069    /// discharged and there is nothing to write.
2070    ///
2071    /// The strongest one is `mfence`, which is what gcc 16.2.0 writes for
2072    /// `__atomic_thread_fence(__ATOMIC_SEQ_CST)` and for `__sync_synchronize`. A locked instruction
2073    /// on the stack is faster on most parts and is what some compilers write instead; it is also a
2074    /// write to memory the program did not ask for, and the plain barrier is the one that says what
2075    /// it means.
2076    ///
2077    /// Written here by name rather than by a rule, for the same reason a `lea` of a symbol is:
2078    /// there is nothing in a barrier that a proof over bitvectors could discharge. It computes
2079    /// nothing, so there is no equality to state, and what makes it the right answer is the memory
2080    /// model, which the rule language cannot talk about.
2081    fn barrier(&mut self, inst: Inst) -> Result<(), Unsupported> {
2082        let Extra::Order(order) = self.source[inst].extra else {
2083            return Err(self.unsupported(inst));
2084        };
2085        if order != MemOrder::SeqCst {
2086            return Ok(());
2087        }
2088        let block = self.at.expect("a block is being filled");
2089        let span = self.source.span(inst);
2090        let fence = mir::Opcode::new(self.names.intern("x64.mfence"));
2091        self.out.build(block, fence).at(span).finish();
2092        Ok(())
2093    }
2094
2095    /// The instruction a program stops on, which is one byte pair and no operands.
2096    ///
2097    /// `ud2` is an opcode the manual promises will never be given a meaning, so a processor that
2098    /// reaches it raises the fault for an instruction it does not know, and on Linux that arrives
2099    /// at the program as `SIGILL`. That is what `__builtin_trap` is for: a stop that cannot be
2100    /// caught by anything the program installed for an ordinary error, cannot be returned from,
2101    /// and leaves the address of the fault in the core file.
2102    ///
2103    /// Why not a call to `abort`. It is two bytes against a call and a relocation, it needs no
2104    /// library, and it works in the places this one is written most, which are a kernel and a
2105    /// freestanding program that has no `abort` to call. gcc 16.2.0 writes `ud2` here too.
2106    fn trap(&mut self, inst: Inst) {
2107        let block = self.at.expect("a block is being filled");
2108        let span = self.source.span(inst);
2109        let stop = mir::Opcode::new(self.names.intern("x64.ud2"));
2110        self.out.build(block, stop).at(span).finish();
2111    }
2112
2113    /// One hint that an address is about to be used, which is one instruction and no promise.
2114    ///
2115    /// Four instructions on this machine and the locality picks between them, which is what the
2116    /// number means: how much of the data will still be wanted after the access. None of it wanted
2117    /// is `prefetchnta`, which brings the line in without keeping it, and all of it wanted is
2118    /// `prefetcht0`, which brings it as close as the machine can. The two in between are the levels
2119    /// between those. Measured against gcc 16.2.0 on x86-64 rather than read off the manual: zero
2120    /// gives `prefetchnta`, one `prefetcht2`, two `prefetcht1` and three `prefetcht0`.
2121    ///
2122    /// Whether the access will write is not read here, and that is this machine rather than an
2123    /// omission. The write hint is `prefetchw`, which is not in the base instruction set, and gcc
2124    /// writes it only when the command line said the part has it. So a prefetch for a write is the
2125    /// same instruction as a prefetch for a read, which is what gcc 16.2.0 writes without
2126    /// `-mprfchw`, and the difference is carried in the IR for a target that can use it.
2127    ///
2128    /// The address goes in the addressing mode rather than in an operand, the way a store's does.
2129    /// It is built here as the plainest one there is, a register and nothing else, because what
2130    /// arrives is a value and folding an addition into the mode is a rule's job and no rule reaches
2131    /// this instruction. An address the program computed is therefore one `lea` or one add in front
2132    /// of this, which is what it would have been for the load the hint is about anyway.
2133    fn hint(&mut self, inst: Inst) -> Result<(), Unsupported> {
2134        let Extra::Prefetch(hint) = self.source[inst].extra else {
2135            return Err(self.unsupported(inst));
2136        };
2137        let args: Vec<Value> = self.source[self.source[inst].args].to_vec();
2138        let [address] = args[..] else { return Err(self.unsupported(inst)) };
2139        let name = match hint.locality {
2140            0 => "prefetch_nta",
2141            1 => "prefetch_t2",
2142            2 => "prefetch_t1",
2143            PrefetchHint::MOST => "prefetch_t0",
2144            // Nothing else exists. The checker reads a locality outside the range as zero and the
2145            // verifier refuses one that got here another way, so this is a hint that was built
2146            // rather than checked, and the safe answer for a hint is to write no instruction.
2147            _ => return Err(self.unsupported(inst)),
2148        };
2149        let base = self.reg_of(address)?;
2150        let block = self.at.expect("a block is being filled");
2151        let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
2152        self.out
2153            .build(block, opcode)
2154            .at(self.source.span(inst))
2155            .mem(mir::Mem::at(mir::Operand::read(base, self.gpr)))
2156            .finish();
2157        Ok(())
2158    }
2159
2160    /// One compare and exchange, which is the instruction every other atomic on this machine is
2161    /// built out of.
2162    ///
2163    /// What the IR asks for is: read what is at an address, compare it against a value the program
2164    /// expected, put a second value there if the two were equal, and say both what was read and
2165    /// whether the exchange happened. The machine has exactly that instruction, and the `lock` in
2166    /// front of it is what makes the whole of it one step as far as every other processor is
2167    /// concerned.
2168    ///
2169    /// The ordering is not read here, and that is the memory model rather than an omission. A
2170    /// locked instruction on x86-64 is a full barrier whatever the program asked for, so a relaxed
2171    /// compare and exchange and a sequentially consistent one are the same instruction, and there
2172    /// is nothing weaker to emit for the weaker orderings. The failure ordering is not read for the
2173    /// same reason.
2174    ///
2175    /// The two values it produces are why this is written by name. The one the program compares
2176    /// against and the one it gets back are both `rax`, which the instruction reads and writes
2177    /// without being told, and the table says so with a fixed constraint at each end rather than
2178    /// leaving the allocator to find out. The second value is the byte behind it, which is the zero
2179    /// flag read out by a `sete`, and it is a definition of the same instruction so that the
2180    /// allocator knows the two are live together and never gives the byte the register the answer
2181    /// is in.
2182    fn exchange(&mut self, inst: Inst) -> Result<(), Unsupported> {
2183        let args: Vec<Value> = self.source[self.source[inst].args].to_vec();
2184        let results: Vec<Value> = self.source[inst].results().collect();
2185        let [addr, expected, desired] = args[..] else { return Err(self.unsupported(inst)) };
2186        let [old, exchanged] = results[..] else { return Err(self.unsupported(inst)) };
2187
2188        // A value the machine can compare in one instruction, which is an integer or an address at
2189        // one of the four widths it has a compare and exchange for. Anything else is a type this
2190        // has no instruction for rather than a program that is wrong, and the front end refuses it
2191        // before ever getting here.
2192        let ty = self.source[old].ty;
2193        let bits = if ty.is_ptr() { ADDRESS_BITS } else { ty.bits() };
2194        if (!ty.is_int() && !ty.is_ptr()) || !matches!(bits, 8 | 16 | 32 | 64) {
2195            return Err(self.unsupported(inst));
2196        }
2197
2198        let base = self.reg_of(addr)?;
2199        let want = self.reg_of(expected)?;
2200        let put = self.reg_of(desired)?;
2201        let got = self.new_reg(old);
2202        let flag = self.new_reg(exchanged);
2203
2204        let name = format!("cmpxchg_{bits}");
2205        let form = x86_64::form(&name).ok_or_else(|| self.unsupported(inst))?;
2206        let block = self.at.expect("a block is being filled");
2207        let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
2208        let mut build = self.out.build(block, opcode).at(self.source.span(inst));
2209        for (desc, reg) in form.operands().iter().zip([got, flag, want, put]) {
2210            let operand = mir::Operand {
2211                reg,
2212                class: desc.class,
2213                role: desc.role,
2214                constraint: desc.constraint,
2215            };
2216            build = build.operand(operand);
2217        }
2218        build.mem(mir::Mem::at(mir::Operand::read(base, self.gpr))).finish();
2219        Ok(())
2220    }
2221
2222    /// One read modify write, for the three operations this machine does in a single instruction.
2223    ///
2224    /// What the IR asks for is: read what is at an address, do something to it, put the answer back,
2225    /// say what was there before, and let nothing get between the three steps. The machine has
2226    /// `xchg` for putting a value there and `lock xadd` for adding one, and both leave what they
2227    /// found in the register the operand arrived in, which is why the value that comes back and the
2228    /// value that went in are one register here.
2229    ///
2230    /// A subtraction is the add over the negated operand, which is right at every width because the
2231    /// machine's arithmetic wraps and negating then adding is subtracting in two's complement
2232    /// whatever the operands were. The negate is a separate instruction in front, over a register of
2233    /// its own, so that the value the program handed over is not the one written on: an operand may
2234    /// be live after this and a program that read it again would read the negation.
2235    ///
2236    /// The ordering is not read, for the reason the compare and exchange beside this does not read
2237    /// it. `xchg` with memory locks the bus whether it is asked to or not and `lock xadd` is asked
2238    /// to, so both are full barriers on this machine and there is nothing weaker to fall to.
2239    ///
2240    /// Eight of the other ten never arrive, because `crate::retry` turned each of them into a loop
2241    /// around a compare and exchange before anything here saw it. The two that do arrive are the
2242    /// ones on floating values, and they are refused: a compare and exchange of a float wants the
2243    /// value carried through an integer of the same width, and an eighty bit float has no such
2244    /// width. Neither family of builtins can write one yet either, so a program that reaches this
2245    /// refusal is a program that reached an unimplemented builtin first.
2246    fn modify(&mut self, inst: Inst) -> Result<(), Unsupported> {
2247        let Extra::Rmw(op, _) = self.source[inst].extra else {
2248            return Err(self.unsupported(inst));
2249        };
2250        let args: Vec<Value> = self.source[self.source[inst].args].to_vec();
2251        let [addr, operand] = args[..] else { return Err(self.unsupported(inst)) };
2252        let old = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
2253
2254        // A value the machine can exchange in one instruction, which is an integer at one of the
2255        // four widths it has these for. A pointer arrives as an address, so it is an integer by the
2256        // time it is here, and anything else is a type this has no instruction for.
2257        let ty = self.source[old].ty;
2258        if !ty.is_int() || !matches!(ty.bits(), 8 | 16 | 32 | 64) {
2259            return Err(self.unsupported(inst));
2260        }
2261        let name = match op {
2262            RmwOp::Xchg => format!("xchg_{}", ty.bits()),
2263            RmwOp::Add | RmwOp::Sub => format!("xadd_{}", ty.bits()),
2264            _ => return Err(self.unsupported(inst)),
2265        };
2266
2267        let base = self.reg_of(addr)?;
2268        let mut put = self.reg_of(operand)?;
2269        let block = self.at.expect("a block is being filled");
2270        let span = self.source.span(inst);
2271        if op == RmwOp::Sub {
2272            let negated = self.out.new_vreg(self.gpr);
2273            let negate =
2274                mir::Opcode::new(self.names.intern(&format!("{PREFIX}neg_r_{}", ty.bits())));
2275            let form = x86_64::form(&format!("neg_r_{}", ty.bits()))
2276                .ok_or_else(|| self.unsupported(inst))?;
2277            let mut build = self.out.build(block, negate).at(span);
2278            for (desc, reg) in form.operands().iter().zip([negated, put]) {
2279                build = build.operand(mir::Operand {
2280                    reg,
2281                    class: desc.class,
2282                    role: desc.role,
2283                    constraint: desc.constraint,
2284                });
2285            }
2286            build.finish();
2287            put = negated;
2288        }
2289
2290        let got = self.new_reg(old);
2291        let form = x86_64::form(&name).ok_or_else(|| self.unsupported(inst))?;
2292        let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
2293        let mut build = self.out.build(block, opcode).at(span);
2294        for (desc, reg) in form.operands().iter().zip([got, put]) {
2295            build = build.operand(mir::Operand {
2296                reg,
2297                class: desc.class,
2298                role: desc.role,
2299                constraint: desc.constraint,
2300            });
2301        }
2302        build.mem(mir::Mem::at(mir::Operand::read(base, self.gpr))).finish();
2303        Ok(())
2304    }
2305
2306    /// One `asm` statement.
2307    ///
2308    /// An empty template is most of the inline assembly in a test suite, and it is not a corner
2309    /// case somebody wrote by accident. A program that wants a value computed where it stands, or a
2310    /// loop the optimizer must not touch, writes `asm volatile ("" : : : "memory")`, and forty
2311    /// years of bug reports about optimizers are full of them. What such a statement asks for is
2312    /// the barrier and the operand places, and no instructions at all.
2313    ///
2314    /// So the operands are the half that is always real: a constraint says where a value has to be,
2315    /// and where it has to be is still true when the template between them is empty.
2316    ///
2317    /// What the constraints ask for, on an empty template, is only ever that two operands share a
2318    /// place. Nothing reads a register no text names, so `"r"` on its own asks for a register and
2319    /// no particular one, and any register at all answers it. A matching constraint is different,
2320    /// because it says the output the assembly leaves is the place the input arrived in, and with
2321    /// no instructions between them that is the input unchanged. So it is a rename and not a move:
2322    /// the value is already in a register and the result is that register.
2323    ///
2324    /// An output nothing is tied to and no instruction writes is whatever the assembly left there,
2325    /// which for a template that writes nothing is whatever was in the register. That is a value
2326    /// the program is not entitled to, and this writes a zero rather than reading one, because the
2327    /// allocator has to be given a definition before a use whatever the program is entitled to.
2328    ///
2329    /// # A template with instructions in it
2330    ///
2331    /// [`x86_64::read`] turns the text into the opcodes this backend already has, which is what
2332    /// `spec/11-asm-objects-debug.md` section 11.1 asks for: the machine is described once, and an
2333    /// instruction a program wrote is looked up in that description rather than copied through to
2334    /// an assembler that has one of its own. So nothing here assembles anything. What it does is
2335    /// put the statement's operands where the opcode holds them, and from there an `asm` statement
2336    /// is ordinary machine code: the allocator picks the registers, the listing and the object file
2337    /// are written from the same table as every other instruction, and a spill around one works
2338    /// because there is nothing left about it for a spill to get wrong.
2339    ///
2340    /// Three things are refused, all for one reason, which is that placing them by a guess gives a
2341    /// program that assembles into something other than what it says.
2342    ///
2343    /// A register the template named itself. The registers an instruction here names are the ones
2344    /// the allocator handed out, and a name in the text is a claim on a register nobody told the
2345    /// allocator about. Doing it properly is the clobber list and a fixed operand, and that is the
2346    /// next piece of this rather than something to approximate now.
2347    ///
2348    /// An output the template writes more than once, and an output that is tied to an input and
2349    /// written. Both are one place with two definitions in it, and the machine IR between here and
2350    /// the allocator has one definition per register by construction.
2351    ///
2352    /// An operand read where the opcode writes, or written where it reads. An output that has not
2353    /// been written yet is not a value, and an input the assembly writes over is a value something
2354    /// else may still be using.
2355    ///
2356    /// The clobber list is still not read. On an empty template that is right rather than an
2357    /// omission, since a template with no instructions in it ruins nothing, and on a template with
2358    /// instructions in it the only registers reachable are the ones named above, which are refused.
2359    fn assembly(&mut self, inst: Inst) -> Result<(), Unsupported> {
2360        let data = &self.source[inst];
2361        let Extra::Asm(asm) = data.extra else { return Err(self.unsupported(inst)) };
2362        let info = self.source[asm];
2363        if !self.source[info.targets].is_empty() {
2364            return Err(Unsupported::Assembly { inst, refused: Written::Goto });
2365        }
2366        let refused = || Unsupported::Assembly { inst, refused: Written::Operand };
2367
2368        let template = self.names.resolve(info.template).to_string();
2369        let lines = if template.trim().is_empty() {
2370            Vec::new()
2371        } else {
2372            x86_64::read(&template)
2373                .ok_or(Unsupported::Assembly { inst, refused: Written::Template })?
2374        };
2375
2376        let constraints = self.names.resolve(info.constraints).to_string();
2377        let results: Vec<Value> = data.results().collect();
2378        let operands = AsmOperands::read(&constraints, &results, &self.source[data.args])
2379            .ok_or_else(refused)?;
2380        let list: Vec<AsmOperand> = operands.iter().copied().collect();
2381
2382        // Which operands the template writes, counted before anything is placed, because the answer
2383        // decides where each of the three below comes from and one instruction may name an operand
2384        // that a later one writes.
2385        let mut writes = vec![0usize; list.len()];
2386        for line in &lines {
2387            let form = x86_64::form(line.opcode).ok_or_else(refused)?;
2388            for (desc, piece) in form.operands().iter().zip(&line.operands) {
2389                let x86_64::Piece::Operand { index, .. } = piece else { continue };
2390                if matches!(desc.role, Role::Def | Role::EarlyDef) {
2391                    *writes.get_mut(*index).ok_or_else(refused)? += 1;
2392                }
2393            }
2394        }
2395
2396        // Where every operand is. Worked out in full before the first instruction is written, since
2397        // reading a value may be what puts it in a register in the first place, and that has to
2398        // happen in front of the assembly rather than in the middle of it.
2399        let mut places: Vec<Option<mir::Reg>> = vec![None; list.len()];
2400        for (index, operand) in list.iter().copied().enumerate() {
2401            let Some(result) = operand.result else {
2402                // An input, or an output the assembly was handed the address of, and both are a
2403                // value that arrives in a register and is read out of it.
2404                places[index] = Some(self.reg_of(operand.value.ok_or_else(refused)?)?);
2405                continue;
2406            };
2407            let ty = self.source[result].ty;
2408            if on_x87(ty) || writes[index] > 1 {
2409                return Err(refused());
2410            }
2411            match operands.tied_to(index) {
2412                // The place the input arrived in, which the assembly wrote nothing over.
2413                Some(from) => {
2414                    if writes[index] > 0 || self.class_of(self.source[from].ty) != self.class_of(ty)
2415                    {
2416                        return Err(refused());
2417                    }
2418                    let reg = self.reg_of(from)?;
2419                    self.regs[result.index()] = Some(reg);
2420                    places[index] = Some(reg);
2421                }
2422                None if writes[index] == 1 => places[index] = Some(self.new_reg(result)),
2423                None => {
2424                    self.undefined(inst, result)?;
2425                    places[index] = self.regs[result.index()];
2426                }
2427            }
2428        }
2429
2430        for line in &lines {
2431            self.instruction(inst, line, &places, &list)?;
2432        }
2433        Ok(())
2434    }
2435
2436    /// One instruction of a template, as the machine instruction it was read back into.
2437    fn instruction(
2438        &mut self,
2439        inst: Inst,
2440        line: &x86_64::Line,
2441        places: &[Option<mir::Reg>],
2442        list: &[AsmOperand],
2443    ) -> Result<(), Unsupported> {
2444        let refused = || Unsupported::Assembly { inst, refused: Written::Operand };
2445        let form = x86_64::form(line.opcode).ok_or_else(refused)?;
2446        let mut built = Vec::with_capacity(line.operands.len());
2447        for (desc, piece) in form.operands().iter().zip(&line.operands) {
2448            built.push(self.placed(inst, *desc, *piece, places, list)?);
2449        }
2450        let at = match line.at {
2451            Some(at) => Some(self.addressed(inst, at, places)?),
2452            None => None,
2453        };
2454
2455        let block = self.at.expect("a block is being filled");
2456        let span = self.source.span(inst);
2457        let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", line.opcode)));
2458        let mut build = self.out.build(block, opcode).at(span);
2459        for operand in built {
2460            build = build.operand(operand);
2461        }
2462        if let Some(value) = line.imm {
2463            build = build.imm(value);
2464        }
2465        if let Some(mem) = at {
2466            build = build.mem(mem);
2467        }
2468        build.finish();
2469        Ok(())
2470    }
2471
2472    /// One operand of one instruction of a template, in the register the statement put it in.
2473    fn placed(
2474        &mut self,
2475        inst: Inst,
2476        desc: OperandDesc,
2477        piece: x86_64::Piece,
2478        places: &[Option<mir::Reg>],
2479        list: &[AsmOperand],
2480    ) -> Result<mir::Operand, Unsupported> {
2481        let refused = || Unsupported::Assembly { inst, refused: Written::Operand };
2482        let x86_64::Piece::Operand { index, width } = piece else { return Err(refused()) };
2483        let operand = list.get(index).copied().ok_or_else(refused)?;
2484        let reg = places.get(index).copied().flatten().ok_or_else(refused)?;
2485
2486        // Read where the opcode reads and written where it writes, which is what the first half of
2487        // this asks. An output has a result and an input has a value, and an output written `+` has
2488        // both, because it is read before it is written.
2489        let placeable = match desc.role {
2490            Role::Use => operand.value.is_some(),
2491            Role::Def | Role::EarlyDef => operand.result.is_some(),
2492        };
2493        let ty = match (operand.result, operand.value) {
2494            (Some(result), _) => self.source[result].ty,
2495            (None, Some(value)) => self.source[value].ty,
2496            (None, None) => return Err(refused()),
2497        };
2498        let bits = if ty.is_ptr() { ADDRESS_BITS } else { ty.bits() };
2499        if !placeable || self.class_of(ty) != desc.class || bits != width.bits() {
2500            return Err(refused());
2501        }
2502        Ok(mir::Operand { reg, class: desc.class, role: desc.role, constraint: desc.constraint })
2503    }
2504
2505    /// The address one instruction of a template reads or writes.
2506    fn addressed(
2507        &mut self,
2508        inst: Inst,
2509        at: x86_64::At,
2510        places: &[Option<mir::Reg>],
2511    ) -> Result<mir::Mem, Unsupported> {
2512        let refused = || Unsupported::Assembly { inst, refused: Written::Operand };
2513        let base = match at.base {
2514            None => None,
2515            Some(x86_64::Piece::Operand { index, .. }) => {
2516                let reg = places.get(index).copied().flatten().ok_or_else(refused)?;
2517                Some(mir::Operand::read(reg, self.gpr))
2518            }
2519            Some(x86_64::Piece::Reg { .. }) => return Err(refused()),
2520        };
2521        Ok(mir::Mem { base, scale: 1, disp: at.disp, segment: at.segment, ..mir::Mem::default() })
2522    }
2523
2524    /// A register holding a value the program has no claim on, written as a zero.
2525    ///
2526    /// Every other way of saying it costs the same instruction or needs a word the machine IR does
2527    /// not have, and a zero is the one that reads the same on every run.
2528    fn undefined(&mut self, inst: Inst, result: Value) -> Result<(), Unsupported> {
2529        let ty = self.source[result].ty;
2530        let refused = Unsupported::Assembly { inst, refused: Written::Operand };
2531        if self.class_of(ty) != self.gpr || !matches!(ty.bits(), 8 | 16 | 32 | 64) {
2532            return Err(refused);
2533        }
2534        let block = self.at.expect("a block is being filled");
2535        let span = self.source.span(inst);
2536        let reg = self.new_reg(result);
2537        let put = mir::Opcode::new(self.names.intern(&format!("{PREFIX}mov_ri_{}", ty.bits())));
2538        self.out.build(block, put).at(span).def(reg, self.gpr).imm(0).finish();
2539        Ok(())
2540    }
2541
2542    /// Whether a type is the width an address is, which is what makes a cast to or from one free.
2543    fn is_address_width(&self, ty: Type) -> bool {
2544        ty.is_ptr() || (ty.is_int() && ty.bits() == ADDRESS_BITS)
2545    }
2546
2547    /// Where a block goes, which in machine IR is on the block rather than on its terminator.
2548    ///
2549    /// That is why no rule ever names a block: a branch is selected for what it reads and the
2550    /// edges are copied across here, arguments and all. The arguments are read last, after every
2551    /// instruction of the block is written, because an argument that is a constant is
2552    /// materialized where it is first wanted and the end of the block is where an edge wants it.
2553    ///
2554    /// Which is not quite the end. A block that leaves two ways has the branch as its last
2555    /// instruction, and a block that leaves through a register has the indirect jump as its last,
2556    /// and anything appended after either is something it has already jumped past, so a constant
2557    /// materialized here would be a register the block below reads and nothing ever writes. The
2558    /// one that was there is put back on the end when that happened, which is the only reordering
2559    /// anything in this crate does and is why it is remembered before a single argument is read.
2560    fn edges(&mut self, block: Block, out: mir::Block) -> Result<(), Unsupported> {
2561        let Some(term) = self.source.terminator(block) else { return Ok(()) };
2562        let leaves = matches!(self.source[term].opcode, Opcode::BrIf | Opcode::IndirectBr);
2563        let branch = if leaves { self.out.terminator(out) } else { None };
2564
2565        let calls: Vec<rucc_ir::BlockCall> = self.source.successors(term).collect();
2566        let mut succs = Vec::with_capacity(calls.len());
2567        for call in calls {
2568            let args: Vec<Value> = self.source[call.args].to_vec();
2569            let mut regs = Vec::with_capacity(args.len());
2570            for value in args {
2571                // The address of where the value is rather than the value, for the one type a
2572                // register holds none of. The block on the other side copies the bytes out of it
2573                // into a slot of its own, which is what makes a second edge into the same block
2574                // safe.
2575                let reg = if on_x87(self.source[value].ty) {
2576                    self.x87_slot(value)
2577                } else {
2578                    self.reg_of(value)?
2579                };
2580                regs.push(reg);
2581            }
2582            succs.push(mir::BlockCall::with(self.out_block(call.block), regs));
2583        }
2584        if let Some(branch) = branch {
2585            if self.out.terminator(out) != Some(branch) {
2586                self.out.remove_inst(branch);
2587                self.out.append_inst(out, branch);
2588            }
2589        }
2590        *self.out.succs_mut(out) = succs;
2591        Ok(())
2592    }
2593
2594    /// The machine IR block an IR block became.
2595    fn out_block(&self, block: Block) -> mir::Block {
2596        self.blocks[block.index()].expect("every block was created before any was filled")
2597    }
2598
2599    /// The parameters of the entry block, which are the function's arguments.
2600    ///
2601    /// They are not block parameters in the machine IR and they cannot be. A block parameter is
2602    /// given its value by a move on the edge into the block, and there is no edge into an entry
2603    /// block, so what arrives in a function is the convention's to say. [`crate::abi`] is what
2604    /// says it.
2605    ///
2606    /// The ones past the last register arrived in the caller's memory and are read out of it, and
2607    /// the loads that read them come back here so that the frame can finish them the way it
2608    /// finishes an `alloca`.
2609    fn arrive(&mut self, block: Block, out: mir::Block) -> Result<(), Unsupported> {
2610        let params = self.source[block].params.clone();
2611        // The type of each is the block's answer and what the ABI asks of it is the signature's,
2612        // and the two lists are the same list: a parameter the classification turned into a
2613        // pointer is a pointer in the block too. A block with more parameters than the signature
2614        // names is not one the front end writes, and each of those is taken as a plain value.
2615        let asked: Vec<Abi> = self.source.signature().params.iter().map(|it| it.abi).collect();
2616        let types: Vec<Param> = params
2617            .iter()
2618            .enumerate()
2619            .map(|(index, &value)| {
2620                let abi = asked.get(index).copied().unwrap_or_default();
2621                Param { ty: self.source[value].ty, abi }
2622            })
2623            .collect();
2624        // A save area for a function that takes arguments its signature does not name, on a
2625        // convention whose list is the four field one. Windows is the other kind and has no area at
2626        // all, so a `va_start` in one is refused rather than built wrong.
2627        let variadic = self.source.signature().variadic && !self.conv.shared_positions;
2628        let area = variadic.then(|| varargs::Area::of(self.conv));
2629        let arrived = abi::entry(&mut self.out, out, &types, self.conv, self.names, area)
2630            .map_err(|(index, missing)| Unsupported::Argument { index, missing })?;
2631        for (&param, reg) in params.iter().zip(&arrived.regs) {
2632            self.regs[param.index()] = Some(*reg);
2633        }
2634        if let Some(area) = area {
2635            self.save_area(out, &arrived, area);
2636        }
2637        self.stack.arguments.extend(arrived.stack);
2638        Ok(())
2639    }
2640
2641    /// The prologue of a variadic function, which is every argument register it was handed written
2642    /// into the frame.
2643    ///
2644    /// Every one the signature did not name, that is. Which of those hold anything is a thing only
2645    /// the caller knew and there is nothing here to ask, so all of them are written, and the ones a
2646    /// named parameter took are not, because `va_start` sets the two offsets past them and nothing
2647    /// ever reads their slots.
2648    ///
2649    /// What that costs is up to fourteen stores in the prologue of a function that may read none of
2650    /// them, and the convention's answer to that is the count of vector registers in `%al`, which
2651    /// lets a callee skip the eight vector stores when the call passed no floats. Skipping them is a
2652    /// branch in a prologue, and a prologue is written long after this by [`crate::finish`], which
2653    /// has no blocks to branch between. So they are all written every time, which is correct and is
2654    /// what `-O0` costs. Issue #323 is the branch.
2655    ///
2656    /// A vector register is written all sixteen bytes at a time, because a `_Float128` fills one and
2657    /// a `va_arg` of a quad reads the slot back whole. gcc writes the same sixteen with the same
2658    /// instruction, which is what [`crate::varargs`] says a list has to be built out of.
2659    ///
2660    /// The address is computed once into a register rather than written as a displacement off the
2661    /// stack pointer, because a displacement into a frame is not known until after allocation and
2662    /// one `lea` costs less than a fixup list for a dozen stores. It is the same `lea` an `alloca`
2663    /// gets and [`crate::finish`] fills it in the same way.
2664    fn save_area(&mut self, out: mir::Block, arrived: &abi::Arrived, area: varargs::Area) {
2665        let save = self.stack.locals.len();
2666        self.stack.locals.push(Local { size: area.size, align: varargs::VECTOR_SLOT });
2667        self.varargs = Some(Varargs {
2668            save,
2669            incoming: arrived.used,
2670            integers: u32::try_from(arrived.took.0).unwrap_or(0) * area.stride(false),
2671            floats: area.starts_at(true)
2672                + u32::try_from(arrived.took.1).unwrap_or(0) * area.stride(true),
2673        });
2674
2675        let base = self.frame_address(out, save);
2676        for &(reg, class, at) in &arrived.spare {
2677            let name = if class == self.gpr { "x64.mov_mr_64" } else { "x64.movaps_mr" };
2678            let store = mir::Opcode::new(self.names.intern(name));
2679            let up = i32::try_from(at).expect("a register save area under two gigabytes");
2680            let mem = mir::Mem::at(mir::Operand::read(base, self.gpr)).plus(up);
2681            self.out.build(out, store).uses(reg, class).mem(mem).finish();
2682        }
2683    }
2684
2685    /// The address of one of the function's stack objects, in a fresh register.
2686    ///
2687    /// Written with nothing in its displacement, because where an object is in a frame is not known
2688    /// until after allocation, and given to [`crate::finish`] to fill in the way an `alloca` is.
2689    fn frame_address(&mut self, out: mir::Block, local: usize) -> mir::Reg {
2690        let reg = self.out.new_vreg(self.gpr);
2691        let lea = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.lea)));
2692        let sp = mir::Operand::read(mir::Reg::physical(self.conv.stack_pointer), self.gpr);
2693        let made = self.out.build(out, lea).def(reg, self.gpr).mem(mir::Mem::at(sp)).finish();
2694        self.stack.addresses.push((made, local));
2695        reg
2696    }
2697
2698    /// Whether an instruction is one no machine instruction is written for where it stands.
2699    ///
2700    /// Four of them, and none is a lowering decision, which is why none is a rule. A constant is
2701    /// written where a register for it is first wanted rather than where the IR put it, and every
2702    /// reader of one may have folded it into an immediate, in which case nowhere is the right
2703    /// place. A return of nothing has nothing to put anywhere: the epilogue gives the frame back
2704    /// and leaves, and it is appended to every block with no successors long after this has
2705    /// finished, so a return with a value is one instruction here and a return without one is
2706    /// none. Unless the value went back through memory, in which case there is something to put
2707    /// somewhere after all and the IR does not carry it: the address the caller handed over has
2708    /// to be in `rax` on the way out, and [`Lowering::returned`] is what writes that.
2709    ///
2710    /// An unconditional jump is the third, and there is even less of it: the edge is on the
2711    /// block, and whether the block it goes to is the next one and needs no jump at all is the
2712    /// block layout's answer rather than this one's.
2713    ///
2714    /// The fourth is a point control does not arrive at, in both of the forms the IR has for it:
2715    /// the `unreachable` terminator the front end puts at the end of a function whose body can run
2716    /// off the bottom, and the `unreachable_hint` a call to `__builtin_unreachable` becomes. What
2717    /// to write for a place nothing reaches is a question with no wrong answer, and nothing is the
2718    /// smallest one and the one gcc 16.2.0 gives at `-O0`. The terminator leaves the block with no
2719    /// successors, so the epilogue lands at the end of it the way it does on any other block that
2720    /// goes nowhere, and the function cannot fall out of its own last instruction into whatever
2721    /// the assembler puts next.
2722    fn writes_nothing(&self, inst: Inst) -> bool {
2723        let data = &self.source[inst];
2724        match data.opcode {
2725            Opcode::IConst | Opcode::Jump | Opcode::Unreachable | Opcode::UnreachableHint => true,
2726            Opcode::Return => self.source[data.args].is_empty() && self.sret().is_none(),
2727            _ => false,
2728        }
2729    }
2730
2731    /// What every instruction in one block matched, with a set of values nobody may take.
2732    ///
2733    /// Backwards, because an instruction that has been folded into a later one does not get to
2734    /// fold anything into itself: the rule that took it only reached one level down, so what is
2735    /// under it is not in the term the matcher saw and cannot be replaced.
2736    fn decide(&self, insts: &[Inst], refused: &HashSet<Value>) -> Decided {
2737        let mut found: Vec<Option<Match<Term>>> = (0..insts.len()).map(|_| None).collect();
2738        let mut plans: Vec<Option<Plan>> = vec![None; insts.len()];
2739        let mut folded: Vec<Inst> = Vec::new();
2740        for (index, &inst) in insts.iter().enumerate().rev() {
2741            if folded.contains(&inst) {
2742                continue;
2743            }
2744            if let Some((plan, matched)) = self.select(inst, refused) {
2745                folded.extend(self.folds(inst, plan));
2746                found[index] = Some(matched);
2747                plans[index] = Some(plan);
2748            }
2749        }
2750        Decided { found, plans, folded }
2751    }
2752
2753    /// A value some of its readers took and some of them did not, which is the one case folding
2754    /// buys nothing.
2755    ///
2756    /// Folding does not delete the instruction that computed a value for anybody else, so a
2757    /// reader that did not take it still needs it in a register and the instruction stays. The
2758    /// reader that did take it now does that work again. Either all of them take it, in which
2759    /// case nothing is left to read it and the instruction goes, or none of them do.
2760    ///
2761    /// The count is over the whole function rather than over the block, since a value read from
2762    /// another block is read from a register there whatever this block decides. An instruction
2763    /// built by name rather than matched, a call being the one that matters, has no plan and so
2764    /// takes nothing, which is the right answer for it as well.
2765    fn left_alive(&self, insts: &[Inst], plans: &[Option<Plan>]) -> Option<Value> {
2766        let mut taken = vec![0u32; self.uses.len()];
2767        for (&inst, plan) in insts.iter().zip(plans) {
2768            let Some(plan) = plan else { continue };
2769            let args = &self.source[self.source[inst].args];
2770            for (index, &arg) in args.iter().take(MAX_ARGS).enumerate() {
2771                if plan[index] == Shown::Expand {
2772                    taken[arg.index()] += 1;
2773                }
2774            }
2775        }
2776        for (&inst, plan) in insts.iter().zip(plans) {
2777            let Some(plan) = plan else { continue };
2778            let args = &self.source[self.source[inst].args];
2779            for (index, &arg) in args.iter().take(MAX_ARGS).enumerate() {
2780                if plan[index] == Shown::Expand && taken[arg.index()] < self.uses[arg.index()] {
2781                    return Some(arg);
2782                }
2783            }
2784        }
2785        None
2786    }
2787
2788    /// The rule that fires on an instruction, and what it bound.
2789    ///
2790    /// The plans are tried in order and the first that matches wins, which is the maximal munch
2791    /// `spec/10-backend.md` asks for: a plan that offers more to the matcher is tried before one
2792    /// that offers less.
2793    fn select(&self, inst: Inst, refused: &HashSet<Value>) -> Option<(Plan, Match<Term>)> {
2794        for plan in self.plans(inst, refused) {
2795            let terms = Terms::new(self.source, inst, plan);
2796            if let Some(matched) = TABLE.find(&terms, Term::Root) {
2797                return Some((plan, matched));
2798            }
2799        }
2800        None
2801    }
2802
2803    /// Every way this instruction can be shown to the matcher, most offered first.
2804    fn plans(&self, inst: Inst, refused: &HashSet<Value>) -> Vec<Plan> {
2805        let args = &self.source[self.source[inst].args];
2806        let mut plans = vec![PLAIN];
2807        for (index, &arg) in args.iter().enumerate().take(MAX_ARGS) {
2808            let mut ways = Vec::new();
2809            if self.foldable(inst, arg, refused) {
2810                ways.push(Shown::Expand);
2811            }
2812            if Terms::new(self.source, inst, PLAIN).constant(arg).is_some() {
2813                ways.push(Shown::Const);
2814            }
2815            ways.push(Shown::Reg);
2816            plans = plans
2817                .into_iter()
2818                .flat_map(|plan| {
2819                    ways.iter().map(move |&way| {
2820                        let mut next = plan;
2821                        next[index] = way;
2822                        next
2823                    })
2824                })
2825                .collect();
2826        }
2827        plans
2828    }
2829
2830    /// Whether an operand may be shown as the instruction that computed it.
2831    ///
2832    /// It has to be in the same block, because a rule that folds one instruction into another
2833    /// moves the work to where the second one is. It has to be something rather than a block
2834    /// parameter, and not a constant, which is shown as a constant instead. And it has to be a
2835    /// value [`Lowering::left_alive`] has not put back, which is how the one reader at a time
2836    /// question is asked here: this says yes to a value with any number of readers, and a value
2837    /// only some of them could take is refused after the fact and asked again.
2838    ///
2839    /// A value with several readers used to be refused outright, on the reasoning that folding
2840    /// does not delete the instruction for anybody else. That reasoning is about the set of
2841    /// readers and was being applied to one reader at a time, which is stricter than it needs to
2842    /// be: when every reader takes it there is nobody left to read it and the instruction goes.
2843    /// An address a store and a load share is the shape that matters, since a memory operand has
2844    /// room for the whole of it and both readers have a memory operand.
2845    fn foldable(&self, into: Inst, value: Value, refused: &HashSet<Value>) -> bool {
2846        let Def::Result { inst, .. } = self.source[value].def else { return false };
2847        if self.source[inst].opcode == Opcode::IConst || refused.contains(&value) {
2848            return false;
2849        }
2850        self.source.block_of(inst).is_some()
2851            && self.source.block_of(inst) == self.source.block_of(into)
2852    }
2853
2854    /// The instructions a match folded into the one it matched.
2855    ///
2856    /// The plan is what says this, not the bindings: a binding is a register or a number either
2857    /// way, and an operand shown as the instruction that computed it is one no rule could have
2858    /// matched without taking that instruction, because the plan offered the matcher nothing
2859    /// else to call it.
2860    fn folds(&self, inst: Inst, plan: Plan) -> Vec<Inst> {
2861        let args = &self.source[self.source[inst].args];
2862        args.iter()
2863            .take(MAX_ARGS)
2864            .enumerate()
2865            .filter(|&(index, _)| plan[index] == Shown::Expand)
2866            .filter_map(|(_, &arg)| match self.source[arg].def {
2867                Def::Result { inst, .. } => Some(inst),
2868                Def::Param { .. } => None,
2869            })
2870            .collect()
2871    }
2872
2873    /// Build the machine instruction a match calls for.
2874    fn emit(&mut self, inst: Inst, matched: &Match<Term>) -> Result<(), Unsupported> {
2875        let rule: &Rule = TABLE.rule(matched);
2876        let pieces = rule.replacement;
2877        let Some(Piece::App { head, arity }) = pieces.first() else {
2878            return Err(self.unsupported(inst));
2879        };
2880        let opcode = head.strip_prefix(PREFIX).ok_or_else(|| self.unsupported(inst))?;
2881        let form = x86_64::form(opcode).ok_or_else(|| self.unsupported(inst))?;
2882
2883        let mut read = Read::default();
2884        let mut at = 1;
2885        for _ in 0..*arity {
2886            at = self.read(inst, pieces, at, &matched.bindings, &mut read)?;
2887        }
2888
2889        let descs = form.operands();
2890        let writes = descs.iter().take_while(|desc| desc.role.is_def()).count();
2891        if descs.len() - writes != read.regs.len() {
2892            return Err(self.unsupported(inst));
2893        }
2894
2895        // The first thing the instruction writes is what it computes, and any others are
2896        // registers the machine destroys on the way, which are fresh because nothing else is in
2897        // them and nothing reads them. An instruction that writes nothing at all is one whose
2898        // whole purpose is its effect, which is what a store is, and there is no result to put
2899        // anywhere.
2900        let mut regs = Vec::new();
2901        if writes > 0 {
2902            let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
2903            regs.push(self.new_reg(result));
2904            // The rest are the registers the machine destroys on the way, and the class each is in
2905            // is the one the instruction's description gives it rather than a guess, so that an
2906            // instruction that wrecks a register in the other file says so.
2907            regs.extend(descs[1..writes].iter().map(|desc| self.out.new_vreg(desc.class)));
2908        } else if self.source[inst].first_result.is_some() {
2909            // A rule that throws away a value the IR gave a name to would leave every reader of
2910            // that name with nothing to read, so it is a rule this and the target disagree about.
2911            return Err(self.unsupported(inst));
2912        }
2913        regs.extend(read.regs.iter().copied());
2914
2915        let block = self.at.expect("a block is being filled");
2916        let opcode = mir::Opcode::new(self.names.intern(head));
2917        let mut build = self.out.build(block, opcode).at(self.source.span(inst));
2918        for (desc, reg) in descs.iter().zip(regs) {
2919            let operand = mir::Operand {
2920                reg,
2921                class: desc.class,
2922                role: desc.role,
2923                constraint: desc.constraint,
2924            };
2925            build = build.operand(operand);
2926        }
2927        if let Some(mem) = read.mem {
2928            build = build.mem(mem);
2929        }
2930        if let Some(imm) = read.imm {
2931            build = build.imm(imm);
2932        }
2933        build.finish();
2934        Ok(())
2935    }
2936
2937    /// Read one argument of a replacement, which is a register, a number or an address.
2938    ///
2939    /// Gives back the position after it, because a replacement is flat and an address takes
2940    /// arguments of its own.
2941    fn read(
2942        &mut self,
2943        inst: Inst,
2944        pieces: &'static [Piece],
2945        at: usize,
2946        bindings: &[Term],
2947        out: &mut Read,
2948    ) -> Result<usize, Unsupported> {
2949        match pieces.get(at) {
2950            Some(Piece::Int(value)) => {
2951                out.imm = i64::try_from(*value).ok();
2952                Ok(at + 1)
2953            }
2954            Some(Piece::Var { index, .. }) => {
2955                match bindings.get(*index) {
2956                    Some(&Term::Reg(value)) => {
2957                        let reg = self.reg_of(value)?;
2958                        out.regs.push(reg);
2959                    }
2960                    Some(&Term::Num(value)) => out.imm = i64::try_from(value).ok(),
2961                    // A pattern binds a register or a number and nothing else, so this is a
2962                    // rule the matcher and this file disagree about.
2963                    _ => return Err(self.unsupported(inst)),
2964                }
2965                Ok(at + 1)
2966            }
2967            Some(Piece::App { head, arity }) => {
2968                let kind = x86_64::address(head).ok_or_else(|| self.unsupported(inst))?;
2969                let mut inner = Read::default();
2970                let mut next = at + 1;
2971                for _ in 0..*arity {
2972                    next = self.read(inst, pieces, next, bindings, &mut inner)?;
2973                }
2974                let mem = address(kind, &inner, self.gpr).ok_or_else(|| self.unsupported(inst))?;
2975                out.mem = Some(mem);
2976                Ok(next)
2977            }
2978            None => Err(self.unsupported(inst)),
2979        }
2980    }
2981
2982    /// The register a value is in, materializing it if it is a constant that has not been put in
2983    /// one yet.
2984    ///
2985    /// A constant is written where it is wanted rather than where the IR defined it, and where it
2986    /// is wanted is a block that need not be the one the IR defined it in. So the register holding
2987    /// one is only good inside the block it was written into, and a second block that wants the
2988    /// same constant gets its own. Anything else is a register read where nothing wrote it: the
2989    /// IR guarantees a definition dominates its uses, and this moved the definition.
2990    ///
2991    /// Writing the number again is also the right answer and not merely the safe one. It is one
2992    /// instruction that reads nothing, which is cheaper than holding a register live across a
2993    /// branch for it, and it is what a rematerializing allocator would do with the value anyway.
2994    fn reg_of(&mut self, value: Value) -> Result<mir::Reg, Unsupported> {
2995        let constant = match self.source[value].def {
2996            Def::Result { inst, .. } => {
2997                (self.source[inst].opcode == Opcode::IConst).then_some(inst)
2998            }
2999            Def::Param { .. } => None,
3000        };
3001        let here = self.at.expect("a block is being filled");
3002        if let Some(reg) = self.regs[value.index()] {
3003            if constant.is_none() || self.written[value.index()] == Some(here) {
3004                return Ok(reg);
3005            }
3006        }
3007        if let Some(inst) = constant {
3008            // Cleared so that the register the constant is written into is a new one rather than
3009            // the one the block above wrote, which is still being read up there.
3010            self.regs[value.index()] = None;
3011            // Nothing is refused here. A constant is written on its own, out of the loop over the
3012            // block, and the operands of the rule that writes one are the number and nothing else.
3013            let matched = self
3014                .select(inst, &HashSet::new())
3015                .map(|(_, matched)| matched)
3016                .ok_or_else(|| self.unsupported(inst))?;
3017            self.emit(inst, &matched)?;
3018            // The same mark the loop over the instructions makes, and it has to be made here as
3019            // well because this is the only place a constant is ever selected: the loop skips one
3020            // where the IR wrote it, so a rule that lowers a constant fires from nowhere else and
3021            // would be reported as a rule nothing reaches.
3022            self.fired.mark(matched.rule);
3023            self.written[value.index()] = Some(here);
3024            return Ok(self.regs[value.index()].expect("a constant is written into a register"));
3025        }
3026        Ok(self.new_reg(value))
3027    }
3028
3029    /// Which register file a value of that type lives in.
3030    ///
3031    /// The vector one for the two float widths the machine has scalar instructions for and for the
3032    /// one it only moves, and the general purpose one for everything else. An eighty bit `long
3033    /// double` is in neither, and it is here rather than in the vector class on purpose: it would
3034    /// be put in a register that cannot hold it, and there is no rule that names one, so the
3035    /// instruction computing it is reported. The wrong class would make that a wrong program
3036    /// instead of a refused one.
3037    ///
3038    /// A hundred and twenty eight bit float is in the vector class and fits it exactly, which is
3039    /// the difference. Nothing computes in it, so every arithmetic on one is still reported, and
3040    /// what the class buys is the moves: a register that holds the whole value is a register a
3041    /// spill, a reload and a copy are each one instruction for.
3042    fn class_of(&self, ty: Type) -> RegClass {
3043        if crate::term::in_vector_file(ty) { self.conv.sse_class } else { self.gpr }
3044    }
3045
3046    /// A fresh register for a value, which is what the instruction computing it writes.
3047    fn new_reg(&mut self, value: Value) -> mir::Reg {
3048        if let Some(reg) = self.regs[value.index()] {
3049            return reg;
3050        }
3051        let reg = self.out.new_vreg(self.class_of(self.source[value].ty));
3052        self.regs[value.index()] = Some(reg);
3053        reg
3054    }
3055
3056    fn unsupported(&self, inst: Inst) -> Unsupported {
3057        let data = &self.source[inst];
3058        Unsupported::Inst {
3059            inst,
3060            term: Terms::new(self.source, inst, PLAIN).name(inst),
3061            opcode: data.opcode,
3062            ty: data.first_result.map(|result| self.source[result].ty),
3063        }
3064    }
3065}
3066
3067/// What the arguments of one replacement came to.
3068#[derive(Debug, Default)]
3069struct Read {
3070    regs: Vec<mir::Reg>,
3071    imm: Option<i64>,
3072    mem: Option<mir::Mem>,
3073}
3074
3075/// The addressing mode an address constructor's arguments make.
3076///
3077/// One arm per constructor rather than a question asked of the kind, because what the arguments
3078/// mean is the whole of what tells the four apart: the same register is a base in one and an
3079/// index in another, and the same constant is a scale in one and a displacement in another.
3080fn address(kind: x86_64::Address, read: &Read, gpr: RegClass) -> Option<mir::Mem> {
3081    let mut regs = read.regs.iter().copied().map(|reg| mir::Operand::read(reg, gpr));
3082    match kind {
3083        x86_64::Address::BaseIndexScale => {
3084            let base = regs.next()?;
3085            let index = regs.next()?;
3086            Some(mir::Mem::at(base).indexed(index, u8::try_from(read.imm?).ok()?))
3087        }
3088        x86_64::Address::IndexScale => Some(mir::Mem {
3089            base: None,
3090            index: Some(regs.next()?),
3091            scale: u8::try_from(read.imm?).ok()?,
3092            disp: 0,
3093            symbol: None,
3094            block: None,
3095            reach: mir::Reach::Itself,
3096            segment: None,
3097        }),
3098        x86_64::Address::Base => Some(mir::Mem::at(regs.next()?)),
3099        // The rule that writes this has a guard saying the constant fits, so a displacement that
3100        // does not is a rule and a target that disagree rather than a program this cannot compile.
3101        x86_64::Address::BaseOffset => {
3102            Some(mir::Mem { disp: i32::try_from(read.imm?).ok()?, ..mir::Mem::at(regs.next()?) })
3103        }
3104    }
3105}
3106
3107/// The table this selector matches with.
3108///
3109/// One target for now, because one target has a rule file. Which table to use becomes a question
3110/// the moment a second one does, and the answer will be the target the session was given rather
3111/// than a constant here.
3112static TABLE: &Table = &crate::select::x86_64::TABLE;
3113
3114#[cfg(test)]
3115mod tests {
3116    use rucc_ir::{
3117        AsmInfo, Builder, CallInfo, Flags, InstData, MemInfo, MemOrder, Restrict, Signature, Type,
3118    };
3119    use rucc_regalloc::assign::Env;
3120    use rucc_target::x86_64::{FRAME, REGS, SYSV};
3121
3122    use super::*;
3123    use crate::finish::{Convention, finish};
3124    use crate::frame::{Frame, Incoming, Layout};
3125
3126    /// A function of as many 64 bit parameters as the test wants, and the block they are in.
3127    fn blank(params: &[Type]) -> (Interner, Func, Block, Vec<Value>) {
3128        let mut names = Interner::new();
3129        let mut func = Func::new(names.intern("f"), Signature::new());
3130        let block = func.create_block();
3131        let values = params.iter().map(|&ty| func.append_param(block, ty)).collect();
3132        (names, func, block, values)
3133    }
3134
3135    /// An ordinary access: not atomic, and aligned enough that nothing here has an opinion.
3136    /// Neither field reaches selection, which is the point of saying it once here.
3137    fn plain() -> MemInfo {
3138        MemInfo {
3139            size: 0,
3140            align: 1,
3141            order: MemOrder::NotAtomic,
3142            tbaa: None,
3143            owns: 0,
3144            restrict: Restrict::NONE,
3145        }
3146    }
3147
3148    /// What the allocator is given: every integer register the convention offers except two, held
3149    /// back so that a move on an edge has somewhere to break a cycle and a spilled value has
3150    /// somewhere to be read into. Which two does not matter, and holding back the last two the
3151    /// convention would reach for leaves every expectation below unchanged.
3152    fn env() -> Env {
3153        const SCRATCH: [rucc_target::PhysReg; 2] = [x86_64::R10, x86_64::R11];
3154        let order: Vec<rucc_target::PhysReg> =
3155            SYSV.int_order.iter().copied().filter(|reg| !SCRATCH.contains(reg)).collect();
3156        Env::new().with(x86_64::GPR, &order, &SCRATCH)
3157    }
3158
3159    /// The machine IR text a function lowers to.
3160    fn lower(names: &mut Interner, source: &Func) -> String {
3161        let out = func(source, names, &SYSV, &Elsewhere::default())
3162            .expect("every instruction has a rule");
3163        mir::print_func(&out.func, names, &REGS)
3164    }
3165
3166    #[test]
3167    fn an_addition_of_two_registers_is_one_instruction() {
3168        let i32 = Type::int(32);
3169        let (mut names, mut func, block, args) = blank(&[i32, i32]);
3170        let mut build = Builder::new(&mut func, block);
3171        build.binary(Opcode::Add, args[0], args[1], Flags::default());
3172
3173        assert_eq!(
3174            lower(&mut names, &func),
3175            "mfunc @f {\nblock0:\n    %0:gpr($rdi) = x64.arg_val_32\n    \
3176             %1:gpr($rsi) = x64.arg_val_32\n    %2:gpr(reuse 1) = x64.add_rr_32 %0, %1\n}\n"
3177        );
3178    }
3179
3180    #[test]
3181    fn a_constant_operand_becomes_an_immediate() {
3182        let i32 = Type::int(32);
3183        let (mut names, mut func, block, args) = blank(&[i32]);
3184        let mut build = Builder::new(&mut func, block);
3185        let seven = build.iconst(i32, 7);
3186        build.binary(Opcode::Add, args[0], seven, Flags::default());
3187
3188        // The constant is in the instruction and nothing was written to hold it, which is what
3189        // materializing one where a register for it is wanted buys.
3190        assert_eq!(
3191            lower(&mut names, &func),
3192            "mfunc @f {\nblock0:\n    %0:gpr($rdi) = x64.arg_val_32\n    \
3193             %1:gpr(reuse 1) = x64.add_ri_32 %0, 7\n}\n"
3194        );
3195    }
3196
3197    #[test]
3198    fn a_constant_too_wide_for_an_immediate_goes_into_a_register() {
3199        let i64 = Type::int(64);
3200        let (mut names, mut func, block, args) = blank(&[i64]);
3201        let mut build = Builder::new(&mut func, block);
3202        let big = build.iconst(i64, i128::from(i32::MAX) + 1);
3203        build.binary(Opcode::Add, args[0], big, Flags::default());
3204
3205        // Nobody wrote this fallback down. The rule that takes an immediate has a guard that
3206        // turns a number this wide down, so it does not fire, and the next way of showing the
3207        // operand puts it in a register.
3208        assert_eq!(
3209            lower(&mut names, &func),
3210            "mfunc @f {\nblock0:\n    %0:gpr($rdi) = x64.arg_val_64\n    \
3211             %1:gpr = x64.mov_ri_64 2147483648\n    %2:gpr(reuse 1) = x64.add_rr_64 %0, %1\n}\n"
3212        );
3213    }
3214
3215    #[test]
3216    fn an_index_calculation_folds_into_an_address() {
3217        let i64 = Type::int(64);
3218        let (mut names, mut func, block, args) = blank(&[i64, i64]);
3219        let mut build = Builder::new(&mut func, block);
3220        let four = build.iconst(i64, 4);
3221        let scaled = build.binary(Opcode::Mul, args[1], four, Flags::default());
3222        build.binary(Opcode::Add, args[0], scaled, Flags::default());
3223
3224        // Three IR instructions and one machine instruction. The multiply is gone because the
3225        // rule that matched reached down and took it.
3226        assert_eq!(
3227            lower(&mut names, &func),
3228            "mfunc @f {\nblock0:\n    %0:gpr($rdi) = x64.arg_val_64\n    \
3229             %1:gpr($rsi) = x64.arg_val_64\n    %2:gpr = x64.lea_64 [%0 + %1*4]\n}\n"
3230        );
3231    }
3232
3233    #[test]
3234    fn an_instruction_every_reader_can_take_is_folded_into_all_of_them() {
3235        let i64 = Type::int(64);
3236        let (mut names, mut func, block, args) = blank(&[i64, i64]);
3237        let mut build = Builder::new(&mut func, block);
3238        let four = build.iconst(i64, 4);
3239        let scaled = build.binary(Opcode::Mul, args[1], four, Flags::default());
3240        let first = build.binary(Opcode::Add, args[0], scaled, Flags::default());
3241        build.binary(Opcode::Add, first, scaled, Flags::default());
3242
3243        // Both readers have room for a scaled index, so both of them take it and nothing is left
3244        // to read the multiply. Three IR instructions become two machine ones, where refusing to
3245        // fold into either reader would have left three.
3246        assert_eq!(
3247            lower(&mut names, &func),
3248            "mfunc @f {\nblock0:\n    %0:gpr($rdi) = x64.arg_val_64\n    \
3249             %1:gpr($rsi) = x64.arg_val_64\n    %2:gpr = x64.lea_64 [%0 + %1*4]\n    \
3250             %3:gpr = x64.lea_64 [%2 + %1*4]\n}\n"
3251        );
3252    }
3253
3254    #[test]
3255    fn an_instruction_one_of_its_readers_cannot_take_is_folded_into_none_of_them() {
3256        let i64 = Type::int(64);
3257        let (mut names, mut func, block, args) = blank(&[i64, i64]);
3258        let mut build = Builder::new(&mut func, block);
3259        let four = build.iconst(i64, 4);
3260        let scaled = build.binary(Opcode::Mul, args[1], four, Flags::default());
3261        build.binary(Opcode::Add, args[0], scaled, Flags::default());
3262        build.store(scaled, args[0], plain(), Flags::default());
3263
3264        // The addition has room for the multiply and the store does not: what a store writes is
3265        // a register, and no rule reaches through it. Folding into the addition alone would
3266        // leave the multiply where it is for the store to read and do the work twice, so the
3267        // multiply is put back and both readers read the register it wrote.
3268        let text = lower(&mut names, &func);
3269        assert!(text.contains("x64.lea_64 [%1*4]"), "{text}");
3270        assert!(text.contains("x64.add_rr_64"), "{text}");
3271    }
3272
3273    #[test]
3274    fn a_shift_by_a_register_asks_for_it_in_cl() {
3275        let i32 = Type::int(32);
3276        let (mut names, mut func, block, args) = blank(&[i32, i32]);
3277        let mut build = Builder::new(&mut func, block);
3278        build.binary(Opcode::Shl, args[0], args[1], Flags::default());
3279
3280        // The fixed register is not in the rule. It is what the target says the instruction does
3281        // with its operands, and the allocator is what will act on it.
3282        let text = lower(&mut names, &func);
3283        assert!(text.contains("x64.shl_rcl_32 %0, %1($rcx)"), "{text}");
3284    }
3285
3286    #[test]
3287    fn a_division_names_the_registers_and_the_register_it_destroys() {
3288        let i32 = Type::int(32);
3289        let (mut names, mut func, block, args) = blank(&[i32, i32]);
3290        let mut build = Builder::new(&mut func, block);
3291        build.binary(Opcode::SDiv, args[0], args[1], Flags::default());
3292
3293        // Two definitions, because a division writes the remainder whether anybody wanted it or
3294        // not, and the second one is early because it is destroyed before the operands are read.
3295        let text = lower(&mut names, &func);
3296        assert!(
3297            text.contains("%2:gpr($rax), early %3:gpr($rdx) = x64.idiv_quo_32 %0($rax), %1"),
3298            "{text}"
3299        );
3300    }
3301
3302    #[test]
3303    fn a_load_reads_through_the_register_the_address_is_in() {
3304        let i64 = Type::int(64);
3305        let (mut names, mut func, block, args) = blank(&[i64]);
3306        let mut build = Builder::new(&mut func, block);
3307        build.load(Type::int(32), args[0], plain(), Flags::default());
3308
3309        assert_eq!(
3310            lower(&mut names, &func),
3311            "mfunc @f {\nblock0:\n    %0:gpr($rdi) = x64.arg_val_64\n    \
3312             %1:gpr = x64.mov_rm_32 [%0]\n}\n"
3313        );
3314    }
3315
3316    #[test]
3317    fn a_store_writes_no_register_and_the_value_it_writes_is_the_one_the_ir_gave_it() {
3318        let (mut names, mut func, block, args) = blank(&[Type::int(32), Type::int(64)]);
3319        let mut build = Builder::new(&mut func, block);
3320        build.store(args[0], args[1], plain(), Flags::default());
3321
3322        // The value is the first parameter and the address is the second, and the instruction
3323        // takes them the other way round. Getting that backwards would compile to a store of the
3324        // address into the value, which is a program that runs and does the wrong thing.
3325        assert_eq!(
3326            lower(&mut names, &func),
3327            "mfunc @f {\nblock0:\n    %0:gpr($rdi) = x64.arg_val_32\n    \
3328             %1:gpr($rsi) = x64.arg_val_64\n    x64.mov_mr_32 %0, [%1]\n}\n"
3329        );
3330    }
3331
3332    #[test]
3333    fn an_address_with_a_constant_added_folds_into_the_access() {
3334        let i64 = Type::int(64);
3335        let (mut names, mut func, block, args) = blank(&[i64]);
3336        let mut build = Builder::new(&mut func, block);
3337        let twelve = build.iconst(i64, 12);
3338        let field = build.binary(Opcode::Add, args[0], twelve, Flags::default());
3339        build.load(Type::int(64), field, plain(), Flags::default());
3340
3341        // Two IR instructions and one machine instruction, which is what every read of a field
3342        // of a structure comes to.
3343        assert_eq!(
3344            lower(&mut names, &func),
3345            "mfunc @f {\nblock0:\n    %0:gpr($rdi) = x64.arg_val_64\n    \
3346             %1:gpr = x64.mov_rm_64 [%0 + 12]\n}\n"
3347        );
3348    }
3349
3350    #[test]
3351    fn a_displacement_too_wide_to_encode_leaves_the_addition_where_it_is() {
3352        let i64 = Type::int(64);
3353        let (mut names, mut func, block, args) = blank(&[i64]);
3354        let mut build = Builder::new(&mut func, block);
3355        let big = build.iconst(i64, i128::from(i32::MAX) + 1);
3356        let far = build.binary(Opcode::Add, args[0], big, Flags::default());
3357        build.load(Type::int(32), far, plain(), Flags::default());
3358
3359        // A displacement is signed and 32 bits. The rule that folds one has a guard that turns
3360        // this down, so the addition stays and the load reads through what it produced. Nobody
3361        // wrote that fallback: it is the next way of showing the operand.
3362        let text = lower(&mut names, &func);
3363        assert!(text.contains("x64.mov_rm_32 [%2]"), "{text}");
3364        assert!(text.contains("x64.add_rr_64"), "{text}");
3365    }
3366
3367    #[test]
3368    fn a_store_of_a_value_that_was_loaded_is_two_instructions_and_no_arithmetic() {
3369        let i64 = Type::int(64);
3370        let (mut names, mut func, block, args) = blank(&[i64, i64]);
3371        let mut build = Builder::new(&mut func, block);
3372        let got = build.load(Type::int(8), args[0], plain(), Flags::default());
3373        build.store(got, args[1], plain(), Flags::default());
3374
3375        // A load feeding a store is the one place folding would be wrong: an x86-64 `mov` has at
3376        // most one memory operand, and there is no rule that takes two, so the load is left where
3377        // it is and the store reads the register it wrote.
3378        assert_eq!(
3379            lower(&mut names, &func),
3380            "mfunc @f {\nblock0:\n    %0:gpr($rdi) = x64.arg_val_64\n    \
3381             %1:gpr($rsi) = x64.arg_val_64\n    %2:gpr = x64.mov_rm_8 [%0]\n    \
3382             x64.mov_mr_8 %2, [%1]\n}\n"
3383        );
3384    }
3385
3386    #[test]
3387    fn an_access_at_a_width_no_rule_is_written_at_is_reported() {
3388        let i64 = Type::int(64);
3389        let (mut names, mut source, block, args) = blank(&[i64]);
3390        let mut build = Builder::new(&mut source, block);
3391        build.load(Type::int(128), args[0], plain(), Flags::default());
3392
3393        // The width is the whole of what is wrong here, so the width is in the message: `load`
3394        // on its own is written about at every other width and would send a reader looking in
3395        // the wrong place.
3396        let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
3397            .expect_err("nothing loads 128 bits");
3398        assert_eq!(failed.to_string(), "no rule lowers a `load` producing a `i128`");
3399    }
3400
3401    #[test]
3402    fn a_return_asks_for_the_value_in_the_register_the_caller_reads() {
3403        let (mut names, mut func, block, args) = blank(&[Type::int(32)]);
3404        let mut build = Builder::new(&mut func, block);
3405        build.ret(&[args[0]]);
3406
3407        // The register is not in the rule, the same way `cl` is not in the rule for a shift. It
3408        // is what the target says the instruction does with its operand, and the allocator is
3409        // what will act on it. There is no `ret` here, because giving the frame back has to
3410        // happen between this and leaving and the frame is not worked out yet.
3411        assert_eq!(
3412            lower(&mut names, &func),
3413            "mfunc @f {\nblock0:\n    %0:gpr($rdi) = x64.arg_val_32\n    \
3414             x64.ret_val_32 %0($rax)\n}\n"
3415        );
3416    }
3417
3418    #[test]
3419    fn a_return_of_two_values_asks_for_the_second_register_as_well() {
3420        let i64 = Type::int(64);
3421        let (mut names, mut func, block, args) = blank(&[i64, i64]);
3422        let mut build = Builder::new(&mut func, block);
3423        build.ret(&[args[0], args[1]]);
3424
3425        // `struct { long a, b; } f(long a, long b)`, after the front end has classified it. Both
3426        // halves are integers, so the second is in the second integer return register, and both
3427        // pseudos say so the same way the one for a single value does.
3428        assert_eq!(
3429            lower(&mut names, &func),
3430            "mfunc @f {\nblock0:\n    %0:gpr($rdi) = x64.arg_val_64\n    \
3431             %1:gpr($rsi) = x64.arg_val_64\n    x64.ret_val_64 %0($rax)\n    \
3432             x64.ret_val2_64 %1($rdx)\n}\n"
3433        );
3434    }
3435
3436    #[test]
3437    fn two_values_back_in_different_files_are_both_the_first_of_their_own() {
3438        let f64 = Type::float(rucc_ir::Float::F64);
3439        let (mut names, mut func, block, args) = blank(&[f64, Type::int(64)]);
3440        let mut build = Builder::new(&mut func, block);
3441        build.ret(&[args[0], args[1]]);
3442
3443        // `struct { double a; long b; } f(double a, long b)`. The two files are counted apart, so
3444        // neither half is the second of anything and the `double` is in `xmm0` rather than in the
3445        // register a second `double` would have been in. Getting this wrong is not a crash: the
3446        // caller reads a register nobody wrote, and this is where that is ruled out.
3447        assert_eq!(
3448            lower(&mut names, &func),
3449            "mfunc @f {\nblock0:\n    %0:xmm($xmm0) = x64.arg_val_f64\n    \
3450             %1:gpr($rdi) = x64.arg_val_64\n    x64.ret_val_f64 %0($xmm0)\n    \
3451             x64.ret_val_64 %1($rax)\n}\n"
3452        );
3453    }
3454
3455    #[test]
3456    fn two_of_the_same_file_back_take_the_first_two_of_it() {
3457        let f64 = Type::float(rucc_ir::Float::F64);
3458        let (mut names, mut func, block, args) = blank(&[f64, f64]);
3459        let mut build = Builder::new(&mut func, block);
3460        build.ret(&[args[0], args[1]]);
3461
3462        // `struct { double x, y; } f(double x, double y)`, which is the vector half of the pair
3463        // above and counts in its own file the same way.
3464        assert_eq!(
3465            lower(&mut names, &func),
3466            "mfunc @f {\nblock0:\n    %0:xmm($xmm0) = x64.arg_val_f64\n    \
3467             %1:xmm($xmm1) = x64.arg_val_f64\n    x64.ret_val_f64 %0($xmm0)\n    \
3468             x64.ret_val2_f64 %1($xmm1)\n}\n"
3469        );
3470    }
3471
3472    /// A function whose answer goes back through memory, with the pointer to the space for it in
3473    /// front of whatever else it takes. Only the signature says it is one.
3474    fn returning_through_memory(params: &[Type]) -> (Interner, Func, Block, Vec<Value>) {
3475        let mut names = Interner::new();
3476        let sret = Abi::Sret { size: 32, align: 8 };
3477        let mut signature = Signature::new().and_param(Param::with_abi(Type::PTR, sret));
3478        signature.params.extend(params.iter().copied().map(Param::new));
3479        let mut func = Func::new(names.intern("f"), signature);
3480        let block = func.create_block();
3481        let space = func.append_param(block, Type::PTR);
3482        let values = std::iter::once(space)
3483            .chain(params.iter().map(|&ty| func.append_param(block, ty)))
3484            .collect();
3485        (names, func, block, values)
3486    }
3487
3488    #[test]
3489    fn the_space_a_return_through_memory_was_given_goes_back_in_the_first_return_register() {
3490        let (mut names, mut func, block, _) = returning_through_memory(&[]);
3491        Builder::new(&mut func, block).ret(&[]);
3492
3493        // `struct big f(void)`, where `big` is too large to come back in registers. The `return`
3494        // carries nothing, because the value went into the space the caller handed over, and the
3495        // document still says that address comes back in `rax`. Nothing in the IR says it, so the
3496        // convention says it, and the pseudo is the one any other pointer return would use.
3497        assert_eq!(
3498            lower(&mut names, &func),
3499            "mfunc @f {\nblock0:\n    %0:gpr($rdi) = x64.arg_val_64\n    \
3500             x64.ret_val_64 %0($rax)\n}\n"
3501        );
3502    }
3503
3504    #[test]
3505    fn what_the_function_did_in_between_does_not_take_the_register_off_it() {
3506        let (mut names, mut func, block, args) = returning_through_memory(&[Type::int(32)]);
3507        let mut build = Builder::new(&mut func, block);
3508        build.store(args[1], args[0], plain(), Flags::default());
3509        build.ret(&[]);
3510
3511        // The register is a read at the end and not a move at the start, so it is live across
3512        // everything between the two and the allocator has to keep it somewhere. In a function
3513        // with a call in it that somewhere is a callee saved register, and the address comes back
3514        // into `rax` here rather than whatever the last instruction happened to leave there. That
3515        // is issue #333, and a store is enough to show the value outlives the entry block.
3516        let text = lower(&mut names, &func);
3517        assert!(text.contains("x64.mov_mr_32 %1, [%0]"), "{text}");
3518        assert!(text.ends_with("    x64.ret_val_64 %0($rax)\n}\n"), "{text}");
3519    }
3520
3521    #[test]
3522    fn a_pointer_that_is_only_a_pointer_is_not_given_back() {
3523        let (mut names, mut func, block, args) = blank(&[Type::PTR]);
3524        let mut build = Builder::new(&mut func, block);
3525        build.store(args[0], args[0], plain(), Flags::default());
3526        build.ret(&[]);
3527
3528        // `void f(void **p)`. It takes a pointer first and returns nothing, which is the shape of
3529        // the one above and none of its meaning, and what tells them apart is the signature. A
3530        // `void` function leaves `rax` alone.
3531        assert!(!lower(&mut names, &func).contains("ret_val"));
3532    }
3533
3534    #[test]
3535    fn a_return_of_a_constant_puts_it_in_a_register_first() {
3536        let (mut names, mut func, block, _) = blank(&[]);
3537        let mut build = Builder::new(&mut func, block);
3538        let zero = build.iconst(Type::int(32), 0);
3539        build.ret(&[zero]);
3540
3541        // No rule returns an immediate, so the plan that offers one is turned down and the next
3542        // one materializes it. That is `int main(void) { return 0; }` in full, once the epilogue
3543        // is appended to it.
3544        assert_eq!(
3545            lower(&mut names, &func),
3546            "mfunc @f {\nblock0:\n    %0:gpr = x64.mov_ri_32 0\n    x64.ret_val_32 %0($rax)\n}\n"
3547        );
3548    }
3549
3550    #[test]
3551    fn the_rule_that_writes_a_constant_down_is_recorded_as_a_rule_that_fired() {
3552        let (mut names, mut func, block, _) = blank(&[]);
3553        let mut build = Builder::new(&mut func, block);
3554        let zero = build.iconst(Type::int(32), 0);
3555        build.ret(&[zero]);
3556
3557        // The loop over the instructions passes a constant by, because a constant is written where
3558        // a register for it is first wanted rather than where the IR put it. So the only place a
3559        // rule about one is ever selected is the materialization, and a mark made in the loop
3560        // alone would report every rule about a constant as a rule nothing reaches.
3561        let out = super::func(&func, &mut names, &SYSV, &Elsewhere::default())
3562            .expect("every instruction has a rule");
3563        let rules = &crate::select::x86_64::TABLE.rules;
3564        let fired: Vec<&str> = rules
3565            .iter()
3566            .enumerate()
3567            .filter(|(index, _)| out.fired.has(*index))
3568            .map(|(_, rule)| rule.pattern)
3569            .collect();
3570        assert!(fired.contains(&"(iconst.i32 k)"), "{fired:?}");
3571    }
3572
3573    #[test]
3574    fn a_return_of_nothing_is_no_instruction_at_all() {
3575        let (mut names, mut func, block, _) = blank(&[]);
3576        let mut build = Builder::new(&mut func, block);
3577        build.ret(&[]);
3578
3579        // Every part of leaving a function that returns nothing is the epilogue's, and the
3580        // epilogue goes in after allocation. A block with nothing in it is the right answer here
3581        // rather than a function that could not be lowered.
3582        assert_eq!(lower(&mut names, &func), "mfunc @f {\nblock0:\n}\n");
3583    }
3584
3585    #[test]
3586    fn the_allocator_is_what_moves_the_answer_into_the_return_register() {
3587        let (mut names, mut source, block, _) = blank(&[]);
3588        let mut build = Builder::new(&mut source, block);
3589        let zero = build.iconst(Type::int(32), 0);
3590        build.ret(&[zero]);
3591
3592        let mut out = func(&source, &mut names, &SYSV, &Elsewhere::default())
3593            .expect("every instruction has a rule")
3594            .func;
3595        let env = env();
3596        let allocation = rucc_regalloc::run(&mut out, &env, "test");
3597        let frame = Frame::of(&out, &allocation, &Layout::new(&SYSV, REGS));
3598        finish(
3599            &mut out,
3600            &allocation,
3601            &frame,
3602            &Stack::default(),
3603            Convention::new(&SYSV, &FRAME),
3604            &mut names,
3605        );
3606
3607        // `int main(void) { return 0; }` end to end. Nothing here asked for `rax`: the rule said
3608        // the value goes back, the target said where, and the allocator is what made it true. The
3609        // epilogue is what leaves, and this function needs no frame, so it is the return alone.
3610        //
3611        // Two instructions and no copy, which is what a hint buys. The return insists on `rax`,
3612        // so `rax` is the register the allocator tries first for the value the return reads, and
3613        // the constant is written straight into it.
3614        assert_eq!(
3615            mir::print_func(&out, &names, &REGS),
3616            "mfunc @f {\nblock0:\n    $rax = x64.mov_ri_32 0\n    \
3617             x64.ret_val_32 $rax($rax)\n    x64.ret\n}\n"
3618        );
3619    }
3620
3621    #[test]
3622    fn a_function_of_two_arguments_is_a_whole_function_now() {
3623        let i32 = Type::int(32);
3624        let (mut names, mut source, block, args) = blank(&[i32, i32]);
3625        let mut build = Builder::new(&mut source, block);
3626        let sum = build.binary(Opcode::Add, args[0], args[1], Flags::default());
3627        build.ret(&[sum]);
3628
3629        let mut out = func(&source, &mut names, &SYSV, &Elsewhere::default())
3630            .expect("every instruction has a rule")
3631            .func;
3632        let env = env();
3633        let allocation = rucc_regalloc::run(&mut out, &env, "test");
3634        let frame = Frame::of(&out, &allocation, &Layout::new(&SYSV, REGS));
3635        finish(
3636            &mut out,
3637            &allocation,
3638            &frame,
3639            &Stack::default(),
3640            Convention::new(&SYSV, &FRAME),
3641            &mut names,
3642        );
3643
3644        // `int f(int a, int b) { return a + b; }` end to end, and this is the test the argument
3645        // side exists for. Before it there was no way to write one: the allocator refuses a
3646        // function whose entry block takes parameters, because there is no edge into an entry
3647        // block for the moves that give a block parameter its value to go on.
3648        //
3649        // One move, and it is the one the machine's addition needs rather than one the allocator
3650        // owes anybody. Each argument stays in the register it arrived in, because the pseudo
3651        // that defines it insists on that register and the allocator now tries it first, and the
3652        // sum stays in the register the addition wrote it to until the return reads it out. The
3653        // copy in front of a two address instruction is what makes its destination one of the
3654        // registers it reads, and the source operand keeps its own name because the destination
3655        // is what the encoder writes.
3656        assert_eq!(
3657            mir::print_func(&out, &names, &REGS),
3658            "mfunc @f {\nblock0:\n    $rdi($rdi) = x64.arg_val_32\n    \
3659             $rsi($rsi) = x64.arg_val_32\n    \
3660             $rdi(reuse 1) = x64.add_rr_32 $rdi, $rsi\n    $rax = x64.mov_rr_64 $rdi\n    \
3661             x64.ret_val_32 $rax($rax)\n    x64.ret\n}\n"
3662        );
3663    }
3664
3665    #[test]
3666    fn an_argument_with_no_register_left_for_it_is_read_out_of_the_caller_s_stack() {
3667        let i64 = Type::int(64);
3668        let (mut names, mut source, block, args) = blank(&[i64; 7]);
3669        let mut build = Builder::new(&mut source, block);
3670        build.ret(&[args[6]]);
3671
3672        let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
3673            .expect("the seventh is read from memory");
3674
3675        // SysV passes six integers in registers and the seventh in the caller's memory, so six of
3676        // these are pseudos that encode to nothing and the seventh is a load that encodes to real
3677        // bytes. Its displacement is nothing here for the reason a local's is: there is no frame
3678        // yet. What the walk hands on is which instruction is waiting, and for how far up the
3679        // caller's argument area, which is the bottom of it because it is the first one there.
3680        assert_eq!(lowered.stack.arguments.len(), 1);
3681        assert_eq!(lowered.stack.arguments[0].1, 0);
3682        let text = mir::print_func(&lowered.func, &names, &REGS);
3683        assert!(text.contains("%6:gpr = x64.mov_rm_64 [$rsp]"), "{text}");
3684        assert_eq!(text.matches("x64.arg_val_64").count(), 6, "{text}");
3685    }
3686
3687    #[test]
3688    fn the_frame_is_what_says_how_far_up_the_caller_s_stack_an_argument_is() {
3689        let i64 = Type::int(64);
3690        let (mut names, mut source, block, args) = blank(&[i64; 8]);
3691        let mut build = Builder::new(&mut source, block);
3692        let sum = build.binary(Opcode::Add, args[6], args[7], Flags::default());
3693        build.ret(&[sum]);
3694
3695        let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
3696            .expect("both are read from memory");
3697        let stack = lowered.stack;
3698        let mut out = lowered.func;
3699        let env = env();
3700        let allocation = rucc_regalloc::run(&mut out, &env, "test");
3701        let layout = stack.layout(Layout::new(&SYSV, REGS));
3702        let frame = Frame::of(&out, &allocation, &layout);
3703        finish(&mut out, &allocation, &frame, &stack, Convention::new(&SYSV, &FRAME), &mut names);
3704
3705        // A leaf that takes no frame, so the stack pointer never moves and the only thing between
3706        // it and the caller's arguments is the return address the call pushed. The seventh
3707        // parameter is at the bottom of the caller's argument area and the eighth is one word
3708        // further up, which is the eight bytes between the two offsets.
3709        let text = mir::print_func(&out, &names, &REGS);
3710        assert_eq!(frame.size(), 0);
3711        assert_eq!(frame.incoming(), Incoming::from_stack(8));
3712        assert!(text.contains("x64.mov_rm_64 [$rsp + 8]"), "{text}");
3713        assert!(text.contains("x64.mov_rm_64 [$rsp + 16]"), "{text}");
3714    }
3715
3716    #[test]
3717    fn a_realigned_frame_reaches_the_caller_s_arguments_through_the_frame_pointer() {
3718        let i64 = Type::int(64);
3719        let (mut names, mut source, block, args) = blank(&[i64; 7]);
3720        let wide = slot(&mut source, block, 64, 32);
3721        let mut build = Builder::new(&mut source, block);
3722        build.store(args[6], wide, plain(), Flags::default());
3723        build.ret(&[args[6]]);
3724
3725        let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
3726            .expect("every instruction has a rule");
3727        let stack = lowered.stack;
3728        let mut out = lowered.func;
3729        let env = env();
3730        let allocation = rucc_regalloc::run(&mut out, &env, "test");
3731        let layout = stack.layout(Layout::new(&SYSV, REGS));
3732        let frame = Frame::of(&out, &allocation, &layout);
3733        finish(&mut out, &allocation, &frame, &stack, Convention::new(&SYSV, &FRAME), &mut names);
3734
3735        // A local wanting thirty two byte alignment makes the prologue force the stack pointer,
3736        // which throws away how far the caller's stack was. So the load the lowering wrote off the
3737        // stack pointer is rewritten to read through the frame pointer, at the one distance that
3738        // survives: the word the prologue pushed the frame pointer into, and the return address
3739        // above it.
3740        let text = mir::print_func(&out, &names, &REGS);
3741        assert_eq!(frame.realign(), Some(32));
3742        assert_eq!(frame.incoming(), Incoming::from_frame(16));
3743        assert!(text.contains("x64.mov_rm_64 [$rbp + 16]"), "{text}");
3744        assert!(!text.contains("x64.mov_rm_64 [$rsp"), "{text}");
3745    }
3746
3747    #[test]
3748    fn a_jump_is_the_edge_and_nothing_else() {
3749        let i32 = Type::int(32);
3750        let (mut names, mut source, entry, args) = blank(&[i32]);
3751        let next = source.create_block();
3752        let got = source.append_param(next, i32);
3753        Builder::new(&mut source, entry).jump(next, &[args[0]]);
3754        Builder::new(&mut source, next).ret(&[got]);
3755
3756        // Two blocks and two instructions, and the jump is neither of them. What it was is the
3757        // arm on the first block, and what the arm carries is the argument it was called with.
3758        assert_eq!(
3759            lower(&mut names, &source),
3760            "mfunc @f {\nblock0:\n    %0:gpr($rdi) = x64.arg_val_32 block1(%0)\n\n\
3761             block1(%1:gpr):\n    x64.ret_val_32 %1($rax)\n}\n"
3762        );
3763    }
3764
3765    /// A block that reads what a block below it writes is filled after it, not before it.
3766    ///
3767    /// The blocks are written entry, `early`, `late`, `exit`, and the entry jumps straight past
3768    /// `early` to `late`, so `late` dominates `early` while sitting below it in the function.
3769    /// Filling them in the order they are written reaches the read in `early` first, and reading
3770    /// a value with no register yet mints one. The cast in `late` is no instruction at all, so
3771    /// what it does is give its answer the register its operand is already in, and that is not
3772    /// the register the read minted. Nothing writes the register the read minted. The printer
3773    /// says `%?` for a register nothing defines, which is what this looks for, and what came out
3774    /// of the real bug was SQLite loading a stack slot no store ever reached.
3775    #[test]
3776    fn a_block_that_reads_what_a_block_below_it_writes_is_filled_after_it() {
3777        let i64 = Type::int(64);
3778        let (mut names, mut source, entry, args) = blank(&[i64, i64]);
3779        let early = source.create_block();
3780        let late = source.create_block();
3781        let exit = source.create_block();
3782
3783        Builder::new(&mut source, entry).jump(late, &[]);
3784        let ptr = cast(&mut source, late, Opcode::IntToPtr, args[0], Type::PTR);
3785        Builder::new(&mut source, early).ret(&[ptr]);
3786        let mut build = Builder::new(&mut source, late);
3787        let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
3788        build.br_if(cond, early, &[], exit, &[]);
3789        Builder::new(&mut source, exit).ret(&[args[1]]);
3790
3791        let text = lower(&mut names, &source);
3792        assert!(!text.contains("%?"), "every register has something that writes it: {text}");
3793    }
3794
3795    /// A constant is written where it is wanted rather than where the IR defined it, and two
3796    /// blocks wanting the same one is two places. Writing it once and reading it in both is a
3797    /// register read where nothing wrote it, unless the block it was written in happens to
3798    /// dominate the other, which nothing here checks and which the second arm of a branch never
3799    /// does. Each block gets its own copy of the number instead.
3800    #[test]
3801    fn a_constant_two_blocks_want_is_written_in_both_of_them() {
3802        let i32 = Type::int(32);
3803        let (mut names, mut source, entry, args) = blank(&[i32, i32]);
3804        let then = source.create_block();
3805        let other = source.create_block();
3806        let join = source.create_block();
3807        let got = source.append_param(join, i32);
3808
3809        let mut build = Builder::new(&mut source, entry);
3810        let seven = build.iconst(i32, 7);
3811        let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
3812        build.br_if(cond, then, &[], other, &[]);
3813        // Both arms want the seven in a register, because a block argument is never an immediate,
3814        // and neither arm dominates the other.
3815        Builder::new(&mut source, then).jump(join, &[seven]);
3816        Builder::new(&mut source, other).jump(join, &[seven]);
3817        Builder::new(&mut source, join).ret(&[got]);
3818
3819        let text = lower(&mut names, &source);
3820        assert_eq!(text.matches("x64.mov_ri_32 7").count(), 2, "one seven per block: {text}");
3821    }
3822
3823    /// An argument on an edge out of a block that leaves two ways is read after every instruction
3824    /// of the block is written, and reading one can write an instruction, which would land after
3825    /// the branch that has already jumped past it. The branch goes back on the end.
3826    #[test]
3827    fn a_constant_an_edge_wants_is_written_before_the_branch_and_not_after_it() {
3828        let i32 = Type::int(32);
3829        let (mut names, mut source, entry, args) = blank(&[i32, i32]);
3830        let then = source.create_block();
3831        let join = source.create_block();
3832        let got = source.append_param(join, i32);
3833
3834        let mut build = Builder::new(&mut source, entry);
3835        let nine = build.iconst(i32, 9);
3836        let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
3837        build.br_if(cond, then, &[], join, &[nine]);
3838        Builder::new(&mut source, then).jump(join, &[args[0]]);
3839        Builder::new(&mut source, join).ret(&[got]);
3840
3841        let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
3842            .expect("every instruction has a rule")
3843            .func;
3844        let entry = out.entry().expect("an entry block");
3845        let last = out.terminator(entry).expect("a block that leaves two ways has a branch");
3846        let branch = names.intern("x64.br_cond_8");
3847        assert_eq!(
3848            out[last].opcode,
3849            mir::Opcode::new(branch),
3850            "the branch is last: {}",
3851            mir::print_func(&out, &names, &REGS)
3852        );
3853    }
3854
3855    #[test]
3856    fn a_conditional_branch_is_lowered_to_the_condition_and_nothing_about_where_it_goes() {
3857        let i32 = Type::int(32);
3858        let (mut names, mut source, entry, args) = blank(&[i32, i32]);
3859        let then = source.create_block();
3860        let other = source.create_block();
3861        let mut build = Builder::new(&mut source, entry);
3862        let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
3863        build.br_if(cond, then, &[], other, &[]);
3864        Builder::new(&mut source, then).ret(&[args[0]]);
3865        Builder::new(&mut source, other).ret(&[args[1]]);
3866
3867        // The comparison writes a byte and the branch reads it, and neither says a block. Both
3868        // arms are on the entry block, in the order the branch took them, so the arm that runs
3869        // when the condition holds is the first.
3870        assert_eq!(
3871            lower(&mut names, &source),
3872            "mfunc @f {\nblock0:\n    %0:gpr($rdi) = x64.arg_val_32\n    \
3873             %1:gpr($rsi) = x64.arg_val_32\n    %2:gpr = x64.cmp_set_l_32 %0, %1\n    \
3874             x64.br_cond_8 %2, block1, block2\n\n\
3875             block1:\n    x64.ret_val_32 %0($rax)\n\n\
3876             block2:\n    x64.ret_val_32 %1($rax)\n}\n"
3877        );
3878    }
3879
3880    /// A choice between two values, which is one instruction and no blocks at all.
3881    ///
3882    /// The arms come out the other way round from the IR, because a conditional move overwrites its
3883    /// destination and the destination is the arm taken when the condition does not hold. The
3884    /// condition arrives last for the same reason: it is read by the test in front of the move
3885    /// rather than by the move.
3886    #[test]
3887    fn a_select_is_lowered_to_a_test_and_a_conditional_move() {
3888        let i32 = Type::int(32);
3889        let (mut names, mut source, entry, args) = blank(&[i32, i32]);
3890        let mut build = Builder::new(&mut source, entry);
3891        let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
3892        let picked = build.select(cond, args[0], args[1]);
3893        build.ret(&[picked]);
3894
3895        assert_eq!(
3896            lower(&mut names, &source),
3897            "mfunc @f {\nblock0:\n    %0:gpr($rdi) = x64.arg_val_32\n    \
3898             %1:gpr($rsi) = x64.arg_val_32\n    %2:gpr = x64.cmp_set_l_32 %0, %1\n    \
3899             %3:gpr(reuse 1) = x64.test_cmov_ne_32 %1, %0, %2\n    \
3900             x64.ret_val_32 %3($rax)\n}\n"
3901        );
3902    }
3903
3904    #[test]
3905    fn a_branch_over_a_block_is_a_whole_function_now() {
3906        let i32 = Type::int(32);
3907        let (mut names, mut source, entry, args) = blank(&[i32, i32]);
3908        let then = source.create_block();
3909        let other = source.create_block();
3910        let join = source.create_block();
3911        let got = source.append_param(join, i32);
3912        let mut build = Builder::new(&mut source, entry);
3913        let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
3914        build.br_if(cond, then, &[], other, &[]);
3915        let mut build = Builder::new(&mut source, then);
3916        let sum = build.binary(Opcode::Add, args[0], args[1], Flags::default());
3917        build.jump(join, &[sum]);
3918        Builder::new(&mut source, other).jump(join, &[args[1]]);
3919        Builder::new(&mut source, join).ret(&[got]);
3920
3921        // `int f(int a, int b) { if (a < b) return a + b; else return b; }` end to end, written
3922        // the way a front end writes it: both arms of the branch are blocks of their own and the
3923        // return is the block they meet at. No edge here is critical, because the two arms out of
3924        // the entry carry nothing and the two arms into the join each leave a block that goes
3925        // nowhere else, so each has its own end to put its move at.
3926        let mut out = func(&source, &mut names, &SYSV, &Elsewhere::default())
3927            .expect("every instruction has a rule")
3928            .func;
3929        assert_eq!(crate::split::critical(&mut out), 0, "no edge here is critical");
3930        let env = env();
3931        let allocation = rucc_regalloc::run(&mut out, &env, "test");
3932        let frame = Frame::of(&out, &allocation, &Layout::new(&SYSV, REGS));
3933        finish(
3934            &mut out,
3935            &allocation,
3936            &frame,
3937            &Stack::default(),
3938            Convention::new(&SYSV, &FRAME),
3939            &mut names,
3940        );
3941
3942        // One epilogue, on the join, which is the one block the function leaves from, and the
3943        // moves that give the join its parameter are at the end of each arm. Every register is
3944        // physical and the branch is still a branch on a register, because turning it into a
3945        // `test` and a `jcc` is the block layout's and there is no block layout yet.
3946        let text = mir::print_func(&out, &names, &REGS);
3947        assert_eq!(text.matches("x64.ret\n").count(), 1, "{text}");
3948        assert!(text.contains("x64.br_cond_8"), "{text}");
3949        assert!(text.contains("x64.add_rr_32"), "{text}");
3950        assert!(!text.contains('%'), "{text}");
3951    }
3952
3953    #[test]
3954    fn a_critical_edge_is_split_before_the_allocator_ever_sees_it() {
3955        let i32 = Type::int(32);
3956        let (mut names, mut source, entry, args) = blank(&[i32, i32]);
3957        let then = source.create_block();
3958        let join = source.create_block();
3959        let got = source.append_param(join, i32);
3960        let mut build = Builder::new(&mut source, entry);
3961        let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
3962        build.br_if(cond, then, &[], join, &[args[1]]);
3963        Builder::new(&mut source, then).jump(join, &[args[0]]);
3964        let mut build = Builder::new(&mut source, join);
3965        let twice = build.binary(Opcode::Add, got, got, Flags::default());
3966        build.ret(&[twice]);
3967
3968        // The else arm is critical: the entry block leaves two ways and the join is arrived at
3969        // two ways, and the arm carries a value. Without splitting it the allocator asserts,
3970        // because the move that gives the join its parameter would have to run at the end of a
3971        // block that also goes to the other arm.
3972        let mut out = func(&source, &mut names, &SYSV, &Elsewhere::default())
3973            .expect("every instruction has a rule")
3974            .func;
3975        assert_eq!(crate::split::critical(&mut out), 1);
3976        let env = env();
3977        let allocation = rucc_regalloc::run(&mut out, &env, "test");
3978        let frame = Frame::of(&out, &allocation, &Layout::new(&SYSV, REGS));
3979        finish(
3980            &mut out,
3981            &allocation,
3982            &frame,
3983            &Stack::default(),
3984            Convention::new(&SYSV, &FRAME),
3985            &mut names,
3986        );
3987
3988        // The block the split added is where the move went, and it is the whole of that block.
3989        let text = mir::print_func(&out, &names, &REGS);
3990        assert_eq!(out.block_count(), 4, "{text}");
3991        assert_eq!(text.matches("x64.ret\n").count(), 1, "{text}");
3992    }
3993
3994    #[test]
3995    fn a_call_passes_what_the_convention_says_and_takes_back_what_it_says() {
3996        let i32 = Type::int(32);
3997        let (mut names, mut source, block, args) = blank(&[i32, i32]);
3998        let sig =
3999            source.add_signature(Signature::new().with_params(&[i32, i32]).with_returns(&[i32]));
4000        let callee = names.intern("g");
4001        let call = Builder::new(&mut source, block).call(callee, sig, &[args[0], args[1]]);
4002        let got = source[call].first_result.expect("an integer comes back");
4003        Builder::new(&mut source, block).ret(&[got]);
4004
4005        // `int f(int a, int b) { return g(a, b); }`. The arguments arrived where the call wants
4006        // them, so what the call reads is what arrived, and the whole of the convention is in the
4007        // constraints rather than in a move.
4008        let text = lower(&mut names, &source);
4009        assert!(text.contains("= x64.call %0($rdi), %1($rsi), @g"), "{text}");
4010        assert!(text.contains("x64.ret_val_32 %2($rax)"), "{text}");
4011        // What the call writes is the value that comes back and then every register the callee is
4012        // free to destroy, in both classes, which is the whole of what stops the allocator from
4013        // leaving something in one of them.
4014        assert!(text.contains("%2:gpr($rax), $rcx, $rdx, $r8, $r9, $r10, $r11, $xmm0,"), "{text}");
4015        assert!(text.contains("$xmm15 = x64.call"), "{text}");
4016    }
4017
4018    #[test]
4019    fn what_the_frame_owes_a_call_comes_back_with_the_function() {
4020        let i32 = Type::int(32);
4021        let sig = |source: &mut Func| source.add_signature(Signature::new().with_params(&[i32]));
4022
4023        let (mut names, mut source, block, args) = blank(&[i32]);
4024        let sig = sig(&mut source);
4025        let callee = names.intern("g");
4026        Builder::new(&mut source, block).call(callee, sig, &[args[0]]);
4027        let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
4028            .expect("every instruction has a rule");
4029
4030        // Nothing on the stack, so nothing owed, but not a leaf either: a function that calls
4031        // owes the callee an aligned stack pointer and may not use the red zone.
4032        assert_eq!(out.stack.calls, Some(0));
4033        let layout = out.stack.layout(Layout::new(&SYSV, REGS));
4034        assert!(!layout.leaf);
4035        assert_eq!(layout.outgoing, 0);
4036
4037        // The same call under the other convention owes thirty two bytes for the callee to spill
4038        // its register arguments into, which is a fact about the convention and not about the call.
4039        let out = func(&source, &mut names, &x86_64::WIN64, &Elsewhere::default())
4040            .expect("every instruction has a rule");
4041        assert_eq!(out.stack.calls, Some(32));
4042
4043        // And a function that calls nothing is a leaf, which is what says it may use the red zone.
4044        let (mut names, mut source, block, args) = blank(&[i32]);
4045        Builder::new(&mut source, block).ret(&[args[0]]);
4046        let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
4047            .expect("every instruction has a rule");
4048        assert_eq!(out.stack.calls, None);
4049        assert!(out.stack.layout(Layout::new(&SYSV, REGS)).leaf);
4050    }
4051
4052    #[test]
4053    fn a_value_that_outlives_a_call_is_not_left_where_the_call_destroys_it() {
4054        let i32 = Type::int(32);
4055        let (mut names, mut source, block, args) = blank(&[i32]);
4056        let sig = source.add_signature(Signature::new().with_params(&[i32]).with_returns(&[i32]));
4057        let callee = names.intern("g");
4058        let call = Builder::new(&mut source, block).call(callee, sig, &[args[0]]);
4059        let got = source[call].first_result.expect("an integer comes back");
4060        let mut build = Builder::new(&mut source, block);
4061        let sum = build.binary(Opcode::Add, got, args[0], Flags::default());
4062        build.ret(&[sum]);
4063
4064        // `int f(int a) { return g(a) + a; }`, which is the smallest program that asks the
4065        // question: `a` is read after the call and `rdi` is a register the call destroys.
4066        let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4067            .expect("every instruction has a rule");
4068        let layout = lowered.stack.layout(Layout::new(&SYSV, REGS));
4069        let mut out = lowered.func;
4070        let env = env();
4071        let allocation = rucc_regalloc::run(&mut out, &env, "test");
4072        let frame = Frame::of(&out, &allocation, &layout);
4073        finish(
4074            &mut out,
4075            &allocation,
4076            &frame,
4077            &Stack::default(),
4078            Convention::new(&SYSV, &FRAME),
4079            &mut names,
4080        );
4081
4082        // It went to a register the callee has to put back, and the prologue and epilogue are what
4083        // put it back, which is the whole bargain the two halves of a convention make.
4084        let text = mir::print_func(&out, &names, &REGS);
4085        assert!(text.contains("$rbx"), "{text}");
4086        assert!(!text.contains('%'), "{text}");
4087        assert_eq!(text.matches("x64.call").count(), 1, "{text}");
4088    }
4089
4090    #[test]
4091    fn a_call_with_more_arguments_than_registers_writes_the_rest_into_the_outgoing_area() {
4092        let i64 = Type::int(64);
4093        let (mut names, mut source, block, args) = blank(&[i64]);
4094        let seven = vec![i64; 7];
4095        let sig = source.add_signature(Signature::new().with_params(&seven));
4096        let callee = names.intern("g");
4097        let passed = vec![args[0]; 7];
4098        Builder::new(&mut source, block).call(callee, sig, &passed);
4099
4100        let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4101            .expect("the seventh goes to memory");
4102        // The bytes the call needs are on the layout the frame is worked out from, so that the
4103        // frame reserves as many as the widest call in the function asked for.
4104        assert_eq!(lowered.stack.calls, Some(8));
4105        let text = mir::print_func(&lowered.func, &names, &REGS);
4106        assert!(text.contains("x64.mov_mr_64 %0, [$rsp]\n"), "{text}");
4107    }
4108
4109    #[test]
4110    fn a_call_this_cannot_make_is_reported_rather_than_made() {
4111        let (mut names, mut source, block, _) = blank(&[]);
4112        let returns = [Type::float(rucc_ir::Float::F80), Type::int(64)];
4113        let sig = source.add_signature(Signature::new().with_returns(&returns));
4114        let callee = names.intern("g");
4115        Builder::new(&mut source, block).call(callee, sig, &[]);
4116        let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4117            .expect_err("a long double is on the x87");
4118        assert_eq!(failed.to_string(), "what this call gives back is on the x87 stack");
4119    }
4120
4121    /// A `long double` on its own is a different answer, because on its own it comes back on the
4122    /// x87 stack rather than in a register, which is somewhere the call cannot be said to write.
4123    ///
4124    /// So the call gives back nothing at all and the value is taken off the stack by the `fstp`
4125    /// straight after it. That instruction has to be straight after it: the stack is one place and
4126    /// anything else that touched it before this ran would be looking at the value still on it.
4127    #[test]
4128    fn a_call_that_gives_back_a_long_double_takes_it_off_the_stack_at_once() {
4129        let (mut names, mut source, block, _) = blank(&[]);
4130        let long_double = Type::float(rucc_ir::Float::F80);
4131        let sig = source.add_signature(Signature::new().with_returns(&[long_double]));
4132        let callee = names.intern("g");
4133        Builder::new(&mut source, block).call(callee, sig, &[]);
4134
4135        let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4136            .expect("the value comes back in st0");
4137        let text = mir::print_func(&lowered.func, &names, &REGS);
4138        let after: Vec<&str> =
4139            text.lines().skip_while(|line| !line.contains("x64.call")).skip(1).collect();
4140        assert_eq!(after[0].trim(), "%0:gpr = x64.lea_64 [$rsp]", "{text}");
4141        assert_eq!(after[1].trim(), "x64.fstp_t [%0]", "{text}");
4142        // And the slot it went into is the sixteen bytes the type takes, like every other one.
4143        assert_eq!(lowered.stack.locals.len(), 1, "{text}");
4144        assert_eq!(lowered.stack.locals[0].size, X87_BYTES);
4145    }
4146
4147    #[test]
4148    fn a_call_through_an_address_goes_through_the_register_the_address_is_in() {
4149        let i32 = Type::int(32);
4150        let (mut names, mut source, block, args) = blank(&[Type::PTR, i32]);
4151        let sig = source.add_signature(Signature::new().with_params(&[i32]).with_returns(&[i32]));
4152        let varargs = source.push_abis(&[]);
4153        let info = source.add_call(CallInfo { callee: None, signature: sig, varargs });
4154        let mut build = Builder::new(&mut source, block);
4155        let inst = InstData {
4156            args: build.func().push_values(&[args[0], args[1]]),
4157            extra: Extra::Call(info),
4158            ..InstData::new(Opcode::CallIndirect)
4159        };
4160        let called = build.inst(inst, &[i32]);
4161        let got = source[called].first_result.expect("an integer comes back");
4162        Builder::new(&mut source, block).ret(&[got]);
4163
4164        // `int f(int (*g)(int), int a) { return g(a); }`. The first operand is the address and
4165        // the arguments are the ones behind it, and everything else about the call is what a call
4166        // to a name would have been.
4167        let text = lower(&mut names, &source);
4168        assert!(text.contains("= x64.call_reg %0, %1($rdi)"), "{text}");
4169        assert!(text.contains("x64.ret_val_32 %2($rax)"), "{text}");
4170        assert!(!text.contains("@g"), "a call through an address names nobody: {text}");
4171    }
4172
4173    #[test]
4174    fn an_instruction_no_rule_covers_is_reported() {
4175        let (mut names, mut source, block, args) = blank(&[Type::PTR]);
4176        let mut build = Builder::new(&mut source, block);
4177        let operands = build.func().push_values(&[args[0]]);
4178        build.inst(InstData { args: operands, ..InstData::new(Opcode::LongjmpMarker) }, &[]);
4179
4180        // The mark that a jump goes back through here, which nothing writes an instruction for
4181        // yet: what it needs is for the allocator to be told a block can be arrived at twice, and
4182        // that is `tamnd/rucc#223`. Nothing about it is a width or a register, so there is nothing
4183        // for the message to add beyond the name.
4184        let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4185            .expect_err("no rule writes a longjmp marker");
4186        assert_eq!(failed.to_string(), "no rule lowers a `longjmp_marker`");
4187
4188        // It produces nothing, so there is no type in the message and nothing invents one, and the
4189        // instruction comes back so a caller can ask the function where it was.
4190        let inst = failed.inst().expect("the instruction it is about");
4191        assert_eq!(source[inst].opcode, Opcode::LongjmpMarker);
4192    }
4193
4194    /// A barrier is written by name here, and what it is depends on the ordering and on nothing
4195    /// else. `crate::expand` is where the reasoning about this machine's memory model lives.
4196    #[test]
4197    fn a_barrier_is_one_instruction_at_the_strongest_ordering_and_none_below_it() {
4198        for order in MemOrder::all().filter(|&order| order != MemOrder::NotAtomic) {
4199            let (mut names, mut source, block, _) = blank(&[]);
4200            let mut build = Builder::new(&mut source, block);
4201            build
4202                .inst(InstData { extra: Extra::Order(order), ..InstData::new(Opcode::Fence) }, &[]);
4203
4204            let text = lower(&mut names, &source);
4205            assert_eq!(text.contains("x64.mfence"), order == MemOrder::SeqCst, "{order:?}: {text}");
4206        }
4207    }
4208
4209    /// A compare and exchange is written by name too, and at the width of the value rather than at
4210    /// the width of the address, which is the mistake worth pinning: everything here is a pointer
4211    /// and only the value says how many bytes the instruction touches.
4212    #[test]
4213    fn a_compare_and_exchange_is_one_instruction_at_the_width_of_the_value() {
4214        for bits in [8, 16, 32, 64] {
4215            let ty = Type::int(bits);
4216            let (mut names, mut source, block, args) = blank(&[Type::PTR, ty, ty]);
4217            let mut build = Builder::new(&mut source, block);
4218            let mem = build.func().add_mem(MemInfo {
4219                size: u64::from(bits / 8),
4220                align: bits / 8,
4221                order: MemOrder::SeqCst,
4222                ..plain()
4223            });
4224            let operands = build.func().push_values(&[args[0], args[1], args[2]]);
4225            build.inst(
4226                InstData {
4227                    args: operands,
4228                    extra: Extra::Mem(mem),
4229                    ..InstData::new(Opcode::Cmpxchg)
4230                },
4231                &[ty, Type::I1],
4232            );
4233
4234            // Two values out of one instruction, the first of them in the register the machine
4235            // reads the expected value out of, the second free for the allocator to place. The
4236            // address is the memory operand and neither of the two values is.
4237            let text = lower(&mut names, &source);
4238            let written = format!("%3:gpr($rax), %4:gpr = x64.cmpxchg_{bits} %1($rax), %2, [%0]");
4239            assert!(text.contains(&written), "{bits}: {text}");
4240        }
4241    }
4242
4243    #[test]
4244    fn more_values_back_than_the_convention_has_registers_for_is_reported() {
4245        let i64 = Type::int(64);
4246        let (mut names, mut source, block, args) = blank(&[i64, i64, i64]);
4247        let mut build = Builder::new(&mut source, block);
4248        build.ret(&[args[0], args[1], args[2]]);
4249
4250        // Two integers come back in `rax` and `rdx` and a third has nowhere to go, which is not a
4251        // gap in the rules but the convention saying no. The front end classifies before it gets
4252        // here, so this is the shape that would mean the classification went wrong.
4253        let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4254            .expect_err("only two come back");
4255        assert_eq!(
4256            failed.to_string(),
4257            "what this function gives back takes more registers than this convention has for it"
4258        );
4259
4260        let inst = failed.inst().expect("the instruction it is about");
4261        assert_eq!(source[inst].opcode, Opcode::Return);
4262    }
4263
4264    /// A refusal about a signature has no instruction, which is what makes it the one arm apart.
4265    ///
4266    /// Everything else is about something written somewhere in the body and hands it back so a
4267    /// caller can ask the function where it came from. A parameter arrives before the first
4268    /// instruction runs, so there is nothing in the body to point at and the message is about
4269    /// the function.
4270    #[test]
4271    fn a_refusal_about_a_parameter_has_no_instruction_to_point_at() {
4272        let missing = Unsupported::Argument { index: 0, missing: Missing::OnX87 };
4273        assert_eq!(missing.inst(), None);
4274    }
4275
4276    /// An `alloca` of a fixed size, which is what every local whose address is taken becomes.
4277    fn slot(source: &mut Func, block: Block, size: u64, align: u32) -> Value {
4278        let info = MemInfo { size, align, ..plain() };
4279        let mut build = Builder::new(source, block);
4280        let mem = build.func().add_mem(info);
4281        build.value(InstData { extra: Extra::Mem(mem), ..InstData::new(Opcode::Alloca) }, Type::PTR)
4282    }
4283
4284    #[test]
4285    fn a_local_is_memory_in_the_frame_and_one_instruction_that_says_where() {
4286        let (mut names, mut source, block, _) = blank(&[]);
4287        let slot = slot(&mut source, block, 4, 4);
4288        let mut build = Builder::new(&mut source, block);
4289        let nine = build.iconst(Type::int(32), 9);
4290        build.store(nine, slot, plain(), Flags::default());
4291        let loaded = build.load(Type::int(32), slot, plain(), Flags::default());
4292        build.ret(&[loaded]);
4293
4294        let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4295            .expect("every instruction has a rule");
4296
4297        // Four bytes on the list the frame is laid out from, and the one instruction that reads
4298        // where they went. Its displacement is nothing here because there is no frame yet, and
4299        // which instruction is waiting for which local is what `finish` is handed.
4300        assert_eq!(lowered.stack.locals, vec![Local { size: 4, align: 4 }]);
4301        assert_eq!(lowered.stack.addresses.len(), 1);
4302        assert_eq!(lowered.stack.addresses[0].1, 0);
4303        assert_eq!(
4304            mir::print_func(&lowered.func, &names, &REGS),
4305            "mfunc @f {\nblock0:\n    %0:gpr = x64.lea_64 [$rsp]\n    \
4306             %1:gpr = x64.mov_ri_32 9\n    x64.mov_mr_32 %1, [%0]\n    \
4307             %2:gpr = x64.mov_rm_32 [%0]\n    x64.ret_val_32 %2($rax)\n}\n"
4308        );
4309    }
4310
4311    #[test]
4312    fn the_frame_is_what_fills_the_address_of_a_local_in() {
4313        let (mut names, mut source, block, _) = blank(&[]);
4314        let slot = slot(&mut source, block, 4, 4);
4315        let mut build = Builder::new(&mut source, block);
4316        let nine = build.iconst(Type::int(32), 9);
4317        build.store(nine, slot, plain(), Flags::default());
4318        let loaded = build.load(Type::int(32), slot, plain(), Flags::default());
4319        build.ret(&[loaded]);
4320
4321        let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4322            .expect("every instruction has a rule");
4323        let stack = lowered.stack;
4324        let mut out = lowered.func;
4325        let env = env();
4326        let allocation = rucc_regalloc::run(&mut out, &env, "test");
4327        let layout = stack.layout(Layout::new(&SYSV, REGS));
4328        let frame = Frame::of(&out, &allocation, &layout);
4329        finish(&mut out, &allocation, &frame, &stack, Convention::new(&SYSV, &FRAME), &mut names);
4330
4331        // `int f(void) { int x; x = 9; return x; }` with the address of `x` taken, end to end.
4332        // A leaf small enough to live in the red zone takes no frame at all, so the stack pointer
4333        // never moves and the four bytes are below it, which is what the negative offset is. The
4334        // instruction the lowering left with nothing in its displacement now has the answer in it.
4335        let text = mir::print_func(&out, &names, &REGS);
4336        assert!(text.contains("$rax = x64.lea_64 [$rsp - 8]"), "{text}");
4337        assert!(!text.contains("x64.sub_ri_64"), "{text}");
4338        assert_eq!(frame.size(), 0);
4339        assert_eq!(frame.local(0), Some(-8));
4340    }
4341
4342    /// An `alloca` whose size is an operand, which is a variable length array.
4343    fn growing(source: &mut Func, block: Block, size: Value, align: u32) -> Value {
4344        let info = MemInfo { size: 0, align, ..plain() };
4345        let mut build = Builder::new(source, block);
4346        let mem = build.func().add_mem(info);
4347        let args = build.func().push_values(&[size]);
4348        build.value(
4349            InstData { args, extra: Extra::Mem(mem), ..InstData::new(Opcode::Alloca) },
4350            Type::PTR,
4351        )
4352    }
4353
4354    #[test]
4355    fn a_stack_slot_whose_size_is_not_known_until_it_runs_takes_the_bytes_off_the_stack_pointer() {
4356        let (mut names, mut source, block, args) = blank(&[Type::int(64)]);
4357        let slot = growing(&mut source, block, args[0], 16);
4358        Builder::new(&mut source, block).ret(&[slot]);
4359
4360        let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4361            .expect("every instruction has a rule");
4362
4363        // The bytes come off the stack pointer where the declaration stands and the address is
4364        // where the stack pointer then is, which is one subtraction and one `lea` rather than a
4365        // slot the frame laid out. Nothing is on the list of locals, because there is nothing
4366        // about this the frame could place.
4367        let text = mir::print_func(&lowered.func, &names, &REGS);
4368        assert!(text.contains("$rsp = x64.sub_rr_64 $rsp, %0"), "{text}");
4369        assert!(text.contains("x64.lea_64 [$rsp]"), "{text}");
4370        assert!(lowered.stack.locals.is_empty(), "{text}");
4371        assert_eq!(lowered.stack.dynamic.len(), 1);
4372        assert!(lowered.stack.grown_at.is_some());
4373    }
4374
4375    #[test]
4376    fn a_growing_slot_wanting_more_alignment_than_the_stack_pointer_has_is_reported() {
4377        let (mut names, mut source, block, args) = blank(&[Type::int(64)]);
4378        let slot = growing(&mut source, block, args[0], 32);
4379        Builder::new(&mut source, block).ret(&[slot]);
4380
4381        // Thirty two is more than a call leaves the stack pointer on, so giving it what it asked
4382        // for means masking the stack pointer after moving it, and after that no constant reaches
4383        // the rest of the frame from the frame pointer either. A second pointer held for the
4384        // purpose is what fixes it and there is not one yet.
4385        let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4386            .expect_err("nothing realigns a frame that grows");
4387        assert_eq!(
4388            failed.to_string(),
4389            "this local wants more alignment than the stack pointer is left on, which needs a \
4390             base register nothing here keeps"
4391        );
4392    }
4393
4394    #[test]
4395    fn a_frame_that_grows_reaches_its_own_locals_through_the_frame_pointer() {
4396        let (mut names, mut source, block, args) = blank(&[Type::int(64)]);
4397        let fixed = slot(&mut source, block, 4, 4);
4398        let mut build = Builder::new(&mut source, block);
4399        let nine = build.iconst(Type::int(32), 9);
4400        build.store(nine, fixed, plain(), Flags::default());
4401        let grown = growing(&mut source, block, args[0], 16);
4402        Builder::new(&mut source, block).ret(&[grown]);
4403
4404        let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4405            .expect("every instruction has a rule");
4406        let stack = lowered.stack;
4407        let mut out = lowered.func;
4408        let env = env();
4409        let allocation = rucc_regalloc::run(&mut out, &env, "test");
4410        let layout = stack.layout(Layout::new(&SYSV, REGS));
4411        let frame = Frame::of(&out, &allocation, &layout);
4412        finish(&mut out, &allocation, &frame, &stack, Convention::new(&SYSV, &FRAME), &mut names);
4413
4414        // The stack pointer moves in the middle of the function, so the four bytes of the fixed
4415        // local are not a constant away from it any more and the frame pointer is what reaches
4416        // them. The frame keeps one whatever the flags asked for, takes its bytes rather than
4417        // living in the red zone, and the address of the growing slot is off the stack pointer as
4418        // it stands after the subtraction rather than off anything the prologue left.
4419        let text = mir::print_func(&out, &names, &REGS);
4420        assert!(frame.grows());
4421        assert!(frame.frame_pointer());
4422        assert!(frame.size() > 0, "{text}");
4423        assert!(text.contains("x64.lea_64 [$rbp"), "{text}");
4424        assert!(text.contains("$rsp = x64.sub_rr_64 $rsp"), "{text}");
4425        assert!(text.contains("x64.lea_64 [$rsp]"), "{text}");
4426    }
4427
4428    #[test]
4429    fn an_address_is_read_written_and_added_to_like_the_integer_it_is() {
4430        let (mut names, mut source, block, args) = blank(&[Type::PTR, Type::int(64)]);
4431        let mut build = Builder::new(&mut source, block);
4432        let stepped = build.func().push_values(&[args[0], args[1]]);
4433        let next =
4434            build.value(InstData { args: stepped, ..InstData::new(Opcode::PtrAdd) }, Type::PTR);
4435        let loaded = build.load(Type::int(32), next, plain(), Flags::default());
4436        build.ret(&[loaded]);
4437
4438        // `int f(int *p, long i) { return *(int *)((char *)p + i); }`. Nothing about this is new
4439        // in the rule set, which is the point: the two addresses arrive in registers because an
4440        // address is an integer as wide as one, and the arithmetic on them is the add it always
4441        // was, so every rule written about an add reaches it.
4442        //
4443        // The add stays its own instruction rather than folding into the address the load reads
4444        // from. Two registers with no scale on either is the one addressing mode the rules have no
4445        // load through, because the folds that exist are the displacement one and the scaled ones,
4446        // and this is neither. That is a peephole worth having and not a thing this changes.
4447        assert_eq!(
4448            lower(&mut names, &source),
4449            "mfunc @f {\nblock0:\n    %0:gpr($rdi) = x64.arg_val_64\n    \
4450             %1:gpr($rsi) = x64.arg_val_64\n    %2:gpr(reuse 1) = x64.add_rr_64 %0, %1\n    \
4451             %3:gpr = x64.mov_rm_32 [%2]\n    x64.ret_val_32 %3($rax)\n}\n"
4452        );
4453    }
4454
4455    /// The address of a file scope name, which is what every use of a global and every string
4456    /// literal starts from.
4457    fn address_of(source: &mut Func, block: Block, names: &mut Interner, name: &str) -> Value {
4458        let symbol = names.intern(name);
4459        let mut build = Builder::new(source, block);
4460        build.value(
4461            InstData { extra: Extra::Symbol(symbol), ..InstData::new(Opcode::GlobalAddr) },
4462            Type::PTR,
4463        )
4464    }
4465
4466    #[test]
4467    fn the_address_of_a_name_is_one_instruction_carrying_the_name() {
4468        let (mut names, mut source, block, _) = blank(&[]);
4469        let counter = address_of(&mut source, block, &mut names, "counter");
4470        let mut build = Builder::new(&mut source, block);
4471        let loaded = build.load(Type::int(32), counter, plain(), Flags::default());
4472        build.ret(&[loaded]);
4473
4474        // `extern int counter; int f(void) { return counter; }`. The address is an addressing mode
4475        // that names no register and carries the symbol, which is what the assembler writes
4476        // relative to `%rip` and what the object writer leaves a relocation for.
4477        assert_eq!(
4478            lower(&mut names, &source),
4479            "mfunc @f {\nblock0:\n    %0:gpr = x64.lea_64 [@counter]\n    \
4480             %1:gpr = x64.mov_rm_32 [%0]\n    x64.ret_val_32 %1($rax)\n}\n"
4481        );
4482    }
4483
4484    #[test]
4485    fn the_address_of_a_name_outside_the_file_is_read_out_of_the_offset_table() {
4486        let (mut names, mut source, block, _) = blank(&[]);
4487        let away = address_of(&mut source, block, &mut names, "away");
4488        Builder::new(&mut source, block).ret(&[away]);
4489        let elsewhere: Elsewhere = [names.intern("away")].into_iter().collect();
4490
4491        // `extern void away(void); void *f(void) { return away; }`. A load and not an address
4492        // computation, because the distance from here to a name a shared library may be the one
4493        // that defines is not a number any link can work out, and the slot the linker fills in is
4494        // in this program and so is a distance it has.
4495        let out =
4496            func(&source, &mut names, &SYSV, &elsewhere).expect("every instruction has a rule");
4497        assert_eq!(
4498            mir::print_func(&out.func, &names, &REGS),
4499            "mfunc @f {\nblock0:\n    %0:gpr = x64.mov_rm_64 [got @away]\n    \
4500             x64.ret_val_64 %0($rax)\n}\n"
4501        );
4502    }
4503
4504    #[test]
4505    fn the_address_of_a_thread_local_is_an_offset_out_of_the_table_plus_where_this_thread_starts() {
4506        let (mut names, mut source, block, _) = blank(&[]);
4507        let own = address_of(&mut source, block, &mut names, "own");
4508        Builder::new(&mut source, block).ret(&[own]);
4509        let elsewhere = Elsewhere::default().with_threads([names.intern("own")]);
4510
4511        // `extern _Thread_local int own; void *f(void) { return &own; }`. Three instructions where
4512        // the two cases above are one, because there is no address to load or to work out: the
4513        // slot holds how far into a thread's block the variable sits, `%fs:0` is where this
4514        // thread's block starts, and the sum of the two is this thread's copy.
4515        let out =
4516            func(&source, &mut names, &SYSV, &elsewhere).expect("every instruction has a rule");
4517        assert_eq!(
4518            mir::print_func(&out.func, &names, &REGS),
4519            "mfunc @f {\nblock0:\n    %0:gpr = x64.mov_rm_64 [thread @own]\n    \
4520             %1:gpr = x64.mov_rm_64 [fs:0]\n    %2:gpr(reuse 1) = x64.add_rr_64 %0, %1\n    \
4521             x64.ret_val_64 %2($rax)\n}\n"
4522        );
4523    }
4524
4525    /// The same load with nothing added to it, which is the whole of `__builtin_thread_pointer`.
4526    #[test]
4527    fn the_start_of_this_thread_s_own_storage_is_the_one_load_and_no_arithmetic() {
4528        let (mut names, mut source, block, _) = blank(&[]);
4529        let here =
4530            Builder::new(&mut source, block).value(InstData::new(Opcode::ThreadPointer), Type::PTR);
4531        Builder::new(&mut source, block).ret(&[here]);
4532
4533        let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
4534            .expect("every instruction has a rule");
4535        assert_eq!(
4536            mir::print_func(&out.func, &names, &REGS),
4537            "mfunc @f {\nblock0:\n    %0:gpr = x64.mov_rm_64 [fs:0]\n    \
4538             x64.ret_val_64 %0($rax)\n}\n"
4539        );
4540    }
4541
4542    /// One `asm` statement, with its template and its constraint list written as a program does.
4543    fn assembly(
4544        source: &mut Func,
4545        block: Block,
4546        names: &mut Interner,
4547        template: &str,
4548        constraints: &str,
4549        args: &[Value],
4550        results: &[Type],
4551    ) -> Inst {
4552        let info = AsmInfo {
4553            template: names.intern(template),
4554            constraints: names.intern(constraints),
4555            clobbers: names.intern("memory"),
4556            targets: rucc_ir::BlockCallList::EMPTY,
4557        };
4558        Builder::new(source, block).inline_asm(info, args, results, Flags::VOLATILE)
4559    }
4560
4561    #[test]
4562    fn an_asm_with_an_empty_template_and_no_operands_is_no_instructions() {
4563        let (mut names, mut source, block, _) = blank(&[]);
4564        assembly(&mut source, block, &mut names, "", "", &[], &[]);
4565        Builder::new(&mut source, block).ret(&[]);
4566
4567        // `asm volatile ("" : : : "memory")`, which is a barrier and nothing else. The barrier was
4568        // spent on the optimizer, which has finished by now, so what is left is nothing.
4569        assert_eq!(lower(&mut names, &source), "mfunc @f {\nblock0:\n}\n");
4570    }
4571
4572    #[test]
4573    fn an_output_an_input_is_tied_to_is_the_register_that_input_arrived_in() {
4574        let i32 = Type::int(32);
4575        let (mut names, mut source, block, args) = blank(&[i32]);
4576        let out = assembly(&mut source, block, &mut names, "", "=r,0", &args, &[i32]);
4577        let produced = source[out].results().next().expect("one result");
4578        Builder::new(&mut source, block).ret(&[produced]);
4579
4580        // `asm ("" : "=r" (x) : "0" (x))`, which is how a program stops the optimizer following a
4581        // value without changing it. The two share a place and the template writes nothing over
4582        // it, so the value comes back out of the register it went in.
4583        assert_eq!(
4584            lower(&mut names, &source),
4585            "mfunc @f {\nblock0:\n    %0:gpr($rdi) = x64.arg_val_32\n    \
4586             x64.ret_val_32 %0($rax)\n}\n"
4587        );
4588    }
4589
4590    #[test]
4591    fn an_output_written_plus_is_the_same_rename() {
4592        let i32 = Type::int(32);
4593        let (mut names, mut source, block, args) = blank(&[i32]);
4594        let out = assembly(&mut source, block, &mut names, "", "+r", &args, &[i32]);
4595        let produced = source[out].results().next().expect("one result");
4596        Builder::new(&mut source, block).ret(&[produced]);
4597
4598        // `asm ("" : "+r" (x))`, which says the same thing in one operand instead of two.
4599        assert_eq!(
4600            lower(&mut names, &source),
4601            "mfunc @f {\nblock0:\n    %0:gpr($rdi) = x64.arg_val_32\n    \
4602             x64.ret_val_32 %0($rax)\n}\n"
4603        );
4604    }
4605
4606    #[test]
4607    fn an_output_nothing_is_tied_to_is_a_zero() {
4608        let i32 = Type::int(32);
4609        let (mut names, mut source, block, _) = blank(&[]);
4610        let out = assembly(&mut source, block, &mut names, "", "=r", &[], &[i32]);
4611        let produced = source[out].results().next().expect("one result");
4612        Builder::new(&mut source, block).ret(&[produced]);
4613
4614        // `asm ("" : "=r" (y))`, whose answer is whatever the assembly left in the register, and
4615        // an empty template leaves nothing. A definite value rather than a register nothing wrote,
4616        // because the allocator is owed a definition before the use however little the program is.
4617        assert_eq!(
4618            lower(&mut names, &source),
4619            "mfunc @f {\nblock0:\n    %0:gpr = x64.mov_ri_32 0\n    x64.ret_val_32 %0($rax)\n}\n"
4620        );
4621    }
4622
4623    #[test]
4624    fn a_template_that_is_one_instruction_becomes_that_instruction() {
4625        let (mut names, mut source, block, _) = blank(&[]);
4626        assembly(&mut source, block, &mut names, "pause", "", &[], &[]);
4627        Builder::new(&mut source, block).ret(&[]);
4628
4629        // `asm volatile ("pause")`, which is what every spin lock in every allocator writes. One
4630        // instruction, no operands, and nothing between the template and the machine but the table
4631        // that already says what a `pause` is.
4632        assert_eq!(lower(&mut names, &source), "mfunc @f {\nblock0:\n    x64.pause\n}\n");
4633    }
4634
4635    #[test]
4636    fn a_template_that_reads_a_segment_becomes_the_load_it_already_was() {
4637        let i64 = Type::int(64);
4638        let (mut names, mut source, block, _) = blank(&[]);
4639        let out = assembly(&mut source, block, &mut names, "movq %%fs:0, %0", "=r", &[], &[i64]);
4640        let produced = source[out].results().next().expect("one result");
4641        Builder::new(&mut source, block).ret(&[produced]);
4642
4643        // `asm ("movq %%fs:0, %0" : "=r" (tid))`, which is how a program finds the block its own
4644        // thread owns. The same instruction `crate::lower` already writes for a thread-local
4645        // variable, reached this time because a program wrote it out by hand.
4646        assert_eq!(
4647            lower(&mut names, &source),
4648            "mfunc @f {\nblock0:\n    %0:gpr = x64.mov_rm_64 [fs:0]\n    \
4649             x64.ret_val_64 %0($rax)\n}\n"
4650        );
4651    }
4652
4653    #[test]
4654    fn a_template_naming_an_instruction_this_machine_has_not_got_is_refused() {
4655        let (mut names, mut source, block, _) = blank(&[]);
4656        assembly(&mut source, block, &mut names, "hcf", "", &[], &[]);
4657        Builder::new(&mut source, block).ret(&[]);
4658
4659        let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4660            .expect_err("there is no such instruction");
4661        assert_eq!(
4662            failed.to_string(),
4663            "this `asm` has instructions in its template, which nothing here assembles"
4664        );
4665    }
4666
4667    /// A register the template named is a claim on a register nobody told the allocator about, and
4668    /// the clobber list that would say so is not read yet. Refused rather than placed, because a
4669    /// register two things believe they own is a wrong program that nothing reports.
4670    #[test]
4671    fn a_template_naming_a_register_the_allocator_did_not_hand_out_is_refused() {
4672        let i64 = Type::int(64);
4673        let (mut names, mut source, block, _) = blank(&[]);
4674        let out = assembly(&mut source, block, &mut names, "movq %%rax, %0", "=r", &[], &[i64]);
4675        let produced = source[out].results().next().expect("one result");
4676        Builder::new(&mut source, block).ret(&[produced]);
4677
4678        let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4679            .expect_err("the template named a register");
4680        assert_eq!(failed.to_string(), "this `asm` has an operand this cannot place");
4681    }
4682
4683    #[test]
4684    fn a_constraint_list_that_does_not_describe_the_operands_is_refused() {
4685        let i32 = Type::int(32);
4686        let (mut names, mut source, block, args) = blank(&[i32]);
4687        assembly(&mut source, block, &mut names, "", "=r", &args, &[]);
4688        Builder::new(&mut source, block).ret(&[]);
4689
4690        // An output with no result to be, which is what the front end never writes and what a
4691        // hand written module can. Refused rather than placed by a guess.
4692        let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4693            .expect_err("the list and the instruction disagree");
4694        assert_eq!(failed.to_string(), "this `asm` has an operand this cannot place");
4695    }
4696
4697    /// A cast between a pointer and an integer, at whatever width the result is asked for.
4698    fn cast(source: &mut Func, block: Block, opcode: Opcode, from: Value, to: Type) -> Value {
4699        let mut build = Builder::new(source, block);
4700        let args = build.func().push_values(&[from]);
4701        build.value(InstData { args, ..InstData::new(opcode) }, to)
4702    }
4703
4704    #[test]
4705    fn a_cast_between_a_pointer_and_an_integer_as_wide_is_no_instruction_at_all() {
4706        let (mut names, mut source, block, args) = blank(&[Type::PTR]);
4707        let number = cast(&mut source, block, Opcode::PtrToInt, args[0], Type::int(64));
4708        Builder::new(&mut source, block).ret(&[number]);
4709
4710        // `long f(void *p) { return (long)p; }`. An address on this machine is an integer as wide
4711        // as the machine addresses, so the cast changes what the type system calls the value and
4712        // changes nothing about the value, and the register holding it is the one that held it.
4713        assert_eq!(
4714            lower(&mut names, &source),
4715            "mfunc @f {\nblock0:\n    %0:gpr($rdi) = x64.arg_val_64\n    \
4716             x64.ret_val_64 %0($rax)\n}\n"
4717        );
4718    }
4719
4720    #[test]
4721    fn a_null_pointer_is_a_constant_that_reaches_a_register_before_anything_reads_it() {
4722        let (mut names, mut source, block, _) = blank(&[]);
4723        let mut build = Builder::new(&mut source, block);
4724        let zero = build.iconst(Type::int(64), 0);
4725        let null = cast(&mut source, block, Opcode::IntToPtr, zero, Type::PTR);
4726        Builder::new(&mut source, block).ret(&[null]);
4727
4728        // `void *f(void) { return 0; }`. The cast is nothing, and reading its operand is what
4729        // writes the zero down: a constant is materialized where it is wanted rather than where
4730        // the IR defined it, and without the read there would be no instruction at all.
4731        assert_eq!(
4732            lower(&mut names, &source),
4733            "mfunc @f {\nblock0:\n    %0:gpr = x64.mov_ri_64 0\n    x64.ret_val_64 %0($rax)\n}\n"
4734        );
4735    }
4736
4737    #[test]
4738    fn the_five_linkages_the_ir_has_narrow_to_the_three_an_object_file_can_say() {
4739        let readings = [
4740            (Linkage::External, mir::Binding::Global),
4741            (Linkage::Common, mir::Binding::Global),
4742            (Linkage::Internal, mir::Binding::Local),
4743            (Linkage::Weak, mir::Binding::Weak),
4744            (Linkage::LinkOnce, mir::Binding::Weak),
4745        ];
4746        for (linkage, wanted) in readings {
4747            let (mut names, mut source, block, _) = blank(&[]);
4748            source.linkage = linkage;
4749            Builder::new(&mut source, block).ret(&[]);
4750            let out = func(&source, &mut names, &SYSV, &Elsewhere::default()).expect("a return");
4751            // The narrowing is done here rather than where the object is written, because a
4752            // machine function is all the assembler and the writer are ever handed.
4753            assert_eq!(out.func.binding, wanted, "{linkage:?}");
4754        }
4755    }
4756
4757    /// The visibility makes the same trip and is not narrowed on the way, because ELF says all
4758    /// three of them.
4759    ///
4760    /// Here for the reason the linkage above is here. A machine function is the whole of what the
4761    /// assembler and the object writer are handed, so a fact about the symbol that does not get
4762    /// onto one is a fact that is gone by the time anything could write it down, and the way that
4763    /// shows up is a shared library exporting the wrong set of names with nothing said anywhere.
4764    #[test]
4765    fn the_visibility_survives_the_trip_from_the_ir_to_a_machine_function() {
4766        let readings = [
4767            (Visibility::Default, mir::Visibility::Default),
4768            (Visibility::Hidden, mir::Visibility::Hidden),
4769            (Visibility::Protected, mir::Visibility::Protected),
4770        ];
4771        for (visibility, wanted) in readings {
4772            let (mut names, mut source, block, _) = blank(&[]);
4773            source.visibility = visibility;
4774            Builder::new(&mut source, block).ret(&[]);
4775            let out = func(&source, &mut names, &SYSV, &Elsewhere::default()).expect("a return");
4776            assert_eq!(out.func.visibility, wanted, "{visibility:?}");
4777        }
4778    }
4779
4780    #[test]
4781    fn a_cast_between_a_pointer_and_a_narrower_integer_is_reported() {
4782        let (mut names, mut source, block, args) = blank(&[Type::PTR]);
4783        let number = cast(&mut source, block, Opcode::PtrToInt, args[0], Type::int(32));
4784        Builder::new(&mut source, block).ret(&[number]);
4785
4786        // The front end never writes one: it casts at the address width and truncates or extends
4787        // around it, so both of those are the rules they always were. IR from somewhere else that
4788        // does write one is refused rather than compiled to a move that keeps the high half.
4789        let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4790            .expect_err("no rule narrows an address");
4791        assert_eq!(failed.to_string(), "no rule lowers a `ptrtoint` producing a `i32`");
4792    }
4793
4794    /// The type this machine has no register for.
4795    fn long_double() -> Type {
4796        Type::float(rucc_ir::Float::F80)
4797    }
4798
4799    #[test]
4800    fn a_double_widened_and_narrowed_again_goes_out_through_the_frame_and_back() {
4801        let f64 = Type::float(rucc_ir::Float::F64);
4802        let (mut names, mut source, block, args) = blank(&[f64]);
4803        let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
4804        let back = cast(&mut source, block, Opcode::FPTrunc, wide, f64);
4805        Builder::new(&mut source, block).ret(&[back]);
4806
4807        // `double f(double d) { long double x = d; return x; }`. The x87 reads memory and nothing
4808        // else, so the value is written to the crossing slot, loaded at the format that widens it
4809        // and put in the slot the eighty bit value lives in. Coming back is the same three the
4810        // other way. Both slots are addressed by a `lea` with nothing in it yet, which is what
4811        // every address in a frame looks like here until `finish` has the numbers.
4812        assert_eq!(
4813            lower(&mut names, &source),
4814            "mfunc @f {\nblock0:\n    \
4815             %0:xmm($xmm0) = x64.arg_val_f64\n    \
4816             %1:gpr = x64.lea_64 [$rsp]\n    \
4817             %2:gpr = x64.lea_64 [$rsp]\n    \
4818             x64.movsd_mr %0, [%1]\n    \
4819             x64.fld_l [%1]\n    \
4820             x64.fstp_t [%2]\n    \
4821             %3:gpr = x64.lea_64 [$rsp]\n    \
4822             %4:gpr = x64.lea_64 [$rsp]\n    \
4823             x64.fld_t [%3]\n    \
4824             x64.fstp_l [%4]\n    \
4825             %5:xmm = x64.movsd_rm [%4]\n    \
4826             x64.ret_val_f64 %5($xmm0)\n}\n"
4827        );
4828    }
4829
4830    #[test]
4831    fn a_long_double_has_sixteen_bytes_of_its_own_and_keeps_them() {
4832        let f64 = Type::float(rucc_ir::Float::F64);
4833        let (mut names, mut source, block, args) = blank(&[f64]);
4834        let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
4835        let once = cast(&mut source, block, Opcode::FPTrunc, wide, f64);
4836        let twice = cast(&mut source, block, Opcode::FPTrunc, wide, f64);
4837        let mut build = Builder::new(&mut source, block);
4838        let sum = build.binary(Opcode::FAdd, once, twice, Flags::default());
4839        build.ret(&[sum]);
4840
4841        let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
4842            .expect("every instruction is written");
4843
4844        // Two slots and not four: sixteen bytes for the one eighty bit value, which is what the
4845        // psABI says one takes and is aligned to, and eight for the crossing, which every group
4846        // in the function shares because nothing is ever left in it. The value's slot is its own
4847        // for the whole function, so reading it twice reads the same sixteen bytes.
4848        assert_eq!(
4849            out.stack.locals,
4850            vec![Local { size: 8, align: 8 }, Local { size: 16, align: 16 }]
4851        );
4852    }
4853
4854    #[test]
4855    fn an_integer_becomes_a_long_double_by_being_loaded_as_one() {
4856        let (mut names, mut source, block, args) = blank(&[Type::int(64)]);
4857        let wide = cast(&mut source, block, Opcode::SIToFP, args[0], long_double());
4858        let back =
4859            cast(&mut source, block, Opcode::FPTrunc, wide, Type::float(rucc_ir::Float::F64));
4860        Builder::new(&mut source, block).ret(&[back]);
4861
4862        // `double f(long n) { long double x = n; return x; }`. `fild` is the same push at another
4863        // format, so the conversion is the load and there is no instruction that converts.
4864        let text = lower(&mut names, &source);
4865        assert!(text.contains("x64.mov_mr_64 %0, [%1]"), "{text}");
4866        assert!(text.contains("x64.fild_ll [%1]"), "{text}");
4867    }
4868
4869    #[test]
4870    fn a_long_double_becoming_an_integer_cuts_towards_zero_with_the_control_word() {
4871        let (mut names, mut source, block, args) = blank(&[Type::float(rucc_ir::Float::F64)]);
4872        let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
4873        let whole = cast(&mut source, block, Opcode::FPToSI, wide, Type::int(32));
4874        Builder::new(&mut source, block).ret(&[whole]);
4875
4876        // The one conversion here with no single instruction behind it. C cuts towards zero and
4877        // the unit rounds the way its control word says, so the word is saved, ORed with the two
4878        // bits that mean truncate, loaded, used and put back. Nine instructions for what `fisttp`
4879        // does in one, and `spec/10-backend.md` section 10.8 says why that one is not used.
4880        let text = lower(&mut names, &source);
4881        let group: Vec<&str> = text
4882            .lines()
4883            .map(str::trim)
4884            .filter(|line| line.starts_with("x64.f") || line.contains("_16"))
4885            .collect();
4886        assert_eq!(
4887            group,
4888            [
4889                "x64.fld_l [%1]",
4890                "x64.fstp_t [%2]",
4891                "x64.fnstcw [%5]",
4892                "%6:gpr = x64.mov_rm_16 [%5]",
4893                "%7:gpr(reuse 1) = x64.or_ri_16 %6, 3072",
4894                "x64.mov_mr_16 %7, [%5 + 2]",
4895                "x64.fldcw [%5 + 2]",
4896                "x64.fld_t [%3]",
4897                "x64.fistp_l [%4]",
4898                "x64.fldcw [%5]",
4899            ],
4900            "{text}"
4901        );
4902    }
4903
4904    #[test]
4905    fn a_long_double_is_read_and_written_as_the_bits_it_already_is() {
4906        let (mut names, mut source, block, args) = blank(&[Type::PTR, Type::PTR]);
4907        let mut build = Builder::new(&mut source, block);
4908        let value = build.load(long_double(), args[0], plain(), Flags::default());
4909        build.store(value, args[1], plain(), Flags::default());
4910        build.ret(&[]);
4911
4912        // `void f(long double *a, long double *b) { *b = *a; }`. A copy is a push and a pop at the
4913        // format the value is already in, which neither converts nor looks: a signalling NaN stays
4914        // one and nothing is raised, which is the whole of what makes it a copy.
4915        let text = lower(&mut names, &source);
4916        let group: Vec<&str> =
4917            text.lines().map(str::trim).filter(|line| line.starts_with("x64.f")).collect();
4918        assert_eq!(
4919            group,
4920            ["x64.fld_t [%0]", "x64.fstp_t [%2]", "x64.fld_t [%3]", "x64.fstp_t [%1]"],
4921            "{text}"
4922        );
4923    }
4924
4925    /// Two `long double` values, from two `double` parameters, and the instructions that made
4926    /// them, which every test below this one throws away.
4927    fn two_long_doubles(source: &mut Func, block: Block, args: &[Value]) -> (Value, Value) {
4928        let left = cast(source, block, Opcode::FPExt, args[0], long_double());
4929        let right = cast(source, block, Opcode::FPExt, args[1], long_double());
4930        (left, right)
4931    }
4932
4933    /// The x87 instructions of a function, in order, with everything else dropped.
4934    fn stack_only(text: &str) -> Vec<&str> {
4935        text.lines().map(str::trim).filter(|line| line.contains("x64.f")).collect()
4936    }
4937
4938    /// The two frame slots the last two addresses of a function were taken of, which in a
4939    /// comparison are the two operands in the order they go on the stack.
4940    fn pushed(out: &Lowered) -> Vec<usize> {
4941        let taken: Vec<usize> = out.stack.addresses.iter().map(|&(_, local)| local).collect();
4942        taken[taken.len() - 2..].to_vec()
4943    }
4944
4945    #[test]
4946    fn adding_two_long_doubles_pushes_both_and_leaves_the_answer_in_a_slot() {
4947        let f64 = Type::float(rucc_ir::Float::F64);
4948        let (mut names, mut source, block, args) = blank(&[f64, f64]);
4949        let (left, right) = two_long_doubles(&mut source, block, &args);
4950        let sum =
4951            Builder::new(&mut source, block).binary(Opcode::FAdd, left, right, Flags::default());
4952        let back = cast(&mut source, block, Opcode::FPTrunc, sum, f64);
4953        Builder::new(&mut source, block).ret(&[back]);
4954
4955        // `double f(double a, double b) { return (long double) a + (long double) b; }`. The last
4956        // four lines are the add: both operands pushed, the instruction that names neither of
4957        // them because they are the top two of a stack, and the answer taken off into its slot.
4958        let text = lower(&mut names, &source);
4959        assert_eq!(
4960            stack_only(&text),
4961            [
4962                "x64.fld_l [%2]",
4963                "x64.fstp_t [%3]",
4964                "x64.fld_l [%4]",
4965                "x64.fstp_t [%5]",
4966                "x64.fld_t [%6]",
4967                "x64.fld_t [%7]",
4968                "x64.fadd_p",
4969                "x64.fstp_t [%8]",
4970                "x64.fld_t [%9]",
4971                "x64.fstp_l [%10]",
4972            ],
4973            "{text}"
4974        );
4975    }
4976
4977    #[test]
4978    fn a_subtraction_pushes_the_left_operand_first_and_asks_for_the_att_spelling() {
4979        let f64 = Type::float(rucc_ir::Float::F64);
4980        let (mut names, mut source, block, args) = blank(&[f64, f64]);
4981        let (left, right) = two_long_doubles(&mut source, block, &args);
4982        let less =
4983            Builder::new(&mut source, block).binary(Opcode::FSub, left, right, Flags::default());
4984        let back = cast(&mut source, block, Opcode::FPTrunc, less, f64);
4985        Builder::new(&mut source, block).ret(&[back]);
4986
4987        // The left one goes on first, so it ends up under the right one, and the answer wanted is
4988        // the one below minus the top. In AT&T that is `fsubrp`, since `fsubp` there is `DE E0+i`
4989        // and computes the other one. The `r` says which spelling this is and not which order the
4990        // pushes were in. `crates/rucc/tests/x87.rs` is what says the answer is right, because a
4991        // name is what got this wrong the first time.
4992        let text = lower(&mut names, &source);
4993        assert_eq!(
4994            &stack_only(&text)[4..8],
4995            ["x64.fld_t [%6]", "x64.fld_t [%7]", "x64.fsubr_p", "x64.fstp_t [%8]"],
4996            "{text}"
4997        );
4998    }
4999
5000    #[test]
5001    fn negating_a_long_double_turns_the_sign_over_and_reads_nothing() {
5002        let f64 = Type::float(rucc_ir::Float::F64);
5003        let (mut names, mut source, block, args) = blank(&[f64]);
5004        let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
5005        let flipped = Builder::new(&mut source, block).unary(Opcode::FNeg, wide, long_double());
5006        let back = cast(&mut source, block, Opcode::FPTrunc, flipped, f64);
5007        Builder::new(&mut source, block).ret(&[back]);
5008
5009        // `fchs` and not a subtraction from zero, which would give a different answer at a negative
5010        // zero and would signal at a NaN. It does not read the value as a number at all.
5011        let text = lower(&mut names, &source);
5012        assert_eq!(
5013            &stack_only(&text)[2..5],
5014            ["x64.fld_t [%3]", "x64.fchs", "x64.fstp_t [%4]"],
5015            "{text}"
5016        );
5017    }
5018
5019    #[test]
5020    fn comparing_two_long_doubles_puts_the_left_one_on_top() {
5021        let f64 = Type::float(rucc_ir::Float::F64);
5022        let (mut names, mut source, block, args) = blank(&[f64, f64]);
5023        let (left, right) = two_long_doubles(&mut source, block, &args);
5024        let mut build = Builder::new(&mut source, block);
5025        build.fcmp(FloatPred::Ogt, left, right, Flags::default());
5026        build.ret(&[]);
5027
5028        // `a > b`. `fucomip` asks about the top of the stack against what is under it, so the
5029        // operand the predicate is about has to go on last, which is the other way round from the
5030        // arithmetic above. The pop that clears the loser and the byte that reads the flags are
5031        // both inside the one opcode.
5032        let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
5033            .expect("every instruction is written");
5034        let slots = pushed(&out);
5035        assert_eq!(slots, [2, 1], "the right operand goes on first and the left one on top");
5036        let text = mir::print_func(&out.func, &names, &REGS);
5037        assert_eq!(
5038            &stack_only(&text)[4..],
5039            ["x64.fld_t [%6]", "x64.fld_t [%7]", "%8:gpr = x64.fucomip_set_a"],
5040            "{text}"
5041        );
5042    }
5043
5044    #[test]
5045    fn a_comparison_that_the_machine_has_backwards_swaps_the_two_pushes() {
5046        let f64 = Type::float(rucc_ir::Float::F64);
5047        let (mut names, mut source, block, args) = blank(&[f64, f64]);
5048        let (left, right) = two_long_doubles(&mut source, block, &args);
5049        let mut build = Builder::new(&mut source, block);
5050        build.fcmp(FloatPred::Olt, left, right, Flags::default());
5051        build.ret(&[]);
5052
5053        // `a < b` is `b > a` and this machine has the one condition, so the same opcode runs with
5054        // the operands the other way round. The same trade the vector rules make, and it has to
5055        // be the same one: a `long double` comparison that picked a different condition from the
5056        // `double` comparison of the same two numbers would be wrong at exactly the unordered
5057        // cases the two conditions differ on.
5058        //
5059        // Which slot each push names is the whole of the difference from the test above, and the
5060        // text does not show it, since an address in a frame is a `lea` with nothing in it until
5061        // `finish` has the numbers. So the slots are what is read here.
5062        let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
5063            .expect("every instruction is written");
5064        let slots = pushed(&out);
5065        assert_eq!(slots, [1, 2], "the left operand goes on first and the right one on top");
5066        let text = mir::print_func(&out.func, &names, &REGS);
5067        assert_eq!(
5068            &stack_only(&text)[4..],
5069            ["x64.fld_t [%6]", "x64.fld_t [%7]", "%8:gpr = x64.fucomip_set_a"],
5070            "{text}"
5071        );
5072    }
5073
5074    #[test]
5075    fn an_ordered_equal_needs_a_second_byte_to_put_the_two_conditions_together() {
5076        let f64 = Type::float(rucc_ir::Float::F64);
5077        let (mut names, mut source, block, args) = blank(&[f64, f64]);
5078        let (left, right) = two_long_doubles(&mut source, block, &args);
5079        let mut build = Builder::new(&mut source, block);
5080        build.fcmp(FloatPred::Oeq, left, right, Flags::default());
5081        build.ret(&[]);
5082
5083        // Equal and ordered are two conditions and the flags carry both, so the opcode writes a
5084        // second register as well as the one the value is in and ANDs them together. Said here by
5085        // handing it a spare, since an instruction that wrote a register nothing knew about would
5086        // be an instruction the allocator could put a live value in the way of.
5087        let text = lower(&mut names, &source);
5088        assert!(text.contains("%8:gpr, %9:gpr = x64.fucomip_set_e_and_np"), "{text}");
5089    }
5090
5091    #[test]
5092    fn a_comparison_that_is_never_asked_is_reported() {
5093        let f64 = Type::float(rucc_ir::Float::F64);
5094        let (mut names, mut source, block, args) = blank(&[f64, f64]);
5095        let (left, right) = two_long_doubles(&mut source, block, &args);
5096        let mut build = Builder::new(&mut source, block);
5097        build.fcmp(FloatPred::False, left, right, Flags::default());
5098        build.ret(&[]);
5099
5100        // Always false is a constant and not a comparison, so there is no condition to pick and
5101        // nothing here folds it into one: an instruction that quietly agreed with it would hide
5102        // that the optimizer left a comparison in that it should have taken out.
5103        let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
5104            .expect_err("no condition is always false");
5105        assert_eq!(failed.to_string(), "no rule lowers a `fcmp` producing a `i1`");
5106    }
5107
5108    #[test]
5109    fn a_long_double_constant_is_the_bits_of_it_put_where_the_value_lives() {
5110        let (mut names, mut source, block, args) = blank(&[Type::PTR]);
5111        let mut build = Builder::new(&mut source, block);
5112        // `1.5L`, which is the leading bit and one more of significand, and an exponent of zero.
5113        let one_and_a_half = build.fconst(long_double(), 0x3fff_c000_0000_0000_0000);
5114        build.store(one_and_a_half, args[0], plain(), Flags::default());
5115        build.ret(&[]);
5116
5117        // No x87 instruction at all. A slot holding one of these is the value, so a constant is
5118        // its ten bytes written where the value lives, and whatever reads it does the `fld`.
5119        let text = lower(&mut names, &source);
5120        assert!(text.contains("x64.mov_ri_64 -4611686018427387904"), "{text}");
5121        assert!(text.contains("x64.mov_ri_16 16383"), "{text}");
5122        assert!(text.contains("x64.mov_mr_16 %3, [%1 + 8]"), "{text}");
5123        // The six bytes above the ten are the padding that makes the type sixteen wide, and they
5124        // are unspecified rather than zero, so nothing writes them.
5125        assert_eq!(text.matches("x64.mov_mr").count(), 2, "{text}");
5126    }
5127
5128    #[test]
5129    fn a_negative_long_double_constant_keeps_the_bit_above_its_exponent() {
5130        let (mut names, mut source, block, args) = blank(&[Type::PTR]);
5131        let mut build = Builder::new(&mut source, block);
5132        let minus = build.fconst(long_double(), 0xbfff_c000_0000_0000_0000);
5133        build.store(minus, args[0], plain(), Flags::default());
5134        build.ret(&[]);
5135
5136        // `-1.5L`. The sign is the top bit of the two byte half, so the immediate that half is put
5137        // in a register with is above the signed range of sixteen bits and has to stay there: read
5138        // as a number it would be negative, and it is not a number, it is two bytes.
5139        let text = lower(&mut names, &source);
5140        assert!(text.contains("x64.mov_ri_16 49151"), "{text}");
5141    }
5142
5143    #[test]
5144    fn a_long_double_crosses_an_edge_as_an_address_and_is_copied_where_it_lands() {
5145        let (mut names, mut source, block, args) = blank(&[Type::float(rucc_ir::Float::F64)]);
5146        let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
5147        let next = source.create_block();
5148        let param = source.append_param(next, long_double());
5149        Builder::new(&mut source, block).jump(next, &[wide]);
5150        Builder::new(&mut source, next).ret(&[param]);
5151
5152        // What the edge carries is the address of the slot the value is already in, which is an
5153        // ordinary register the allocator has an opinion about. The block on the other side copies
5154        // the sixteen bytes into a slot of its own before anything reads them, so a second edge
5155        // handing over a second address would still leave one place for a reader to look.
5156        let text = lower(&mut names, &source);
5157        let second: Vec<&str> = text
5158            .lines()
5159            .skip_while(|line| !line.starts_with("block1"))
5160            .skip(1)
5161            .take(3)
5162            .map(str::trim)
5163            .collect();
5164        assert_eq!(
5165            second,
5166            ["x64.fld_t [%4]", "%5:gpr = x64.lea_64 [$rsp]", "x64.fstp_t [%5]"],
5167            "{text}"
5168        );
5169    }
5170
5171    #[test]
5172    fn more_long_doubles_at_a_block_than_the_stack_is_deep_are_reported() {
5173        let f64 = Type::float(rucc_ir::Float::F64);
5174        let (mut names, mut source, block, args) = blank(&[f64]);
5175        let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
5176        let next = source.create_block();
5177        let params: Vec<Value> =
5178            (0..=X87_DEPTH).map(|_| source.append_param(next, long_double())).collect();
5179        let carried: Vec<Value> = params.iter().map(|_| wide).collect();
5180        Builder::new(&mut source, block).jump(next, &carried);
5181        Builder::new(&mut source, next).ret(&[params[0]]);
5182
5183        // The copies go through the x87 stack so that every one of them is read before any of them
5184        // is written, which is what makes a block that swaps two of these right. Nine of them do
5185        // not fit on the stack, and copying the ninth before or after the rest is the order that
5186        // could be wrong, so it is refused instead.
5187        let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
5188            .expect_err("nine do not fit on the stack");
5189        assert_eq!(
5190            failed.to_string(),
5191            "block1 takes 9 parameters of type `f80` and only 8 can cross an edge at once"
5192        );
5193        assert_eq!(failed.inst(), None);
5194    }
5195}