rucc_codegen/lower.rs
1//! The selector: an IR function becomes a machine IR function.
2//!
3//! Design: `spec/10-backend.md` sections 10.2 and 10.3.
4//!
5//! What the matcher in [`crate::select`] does is answer one question about one term. What this
6//! does is ask it: walk a function, decide which terms are worth asking about, and build machine
7//! instructions out of what comes back. Nothing here decides what an IR term lowers to. That is
8//! in `rules/x86-64.rules` and it is proved before it is used, which is the whole point of the
9//! arrangement and the reason this file is short.
10//!
11//! # What it does with an instruction
12//!
13//! It tries the ways the instruction can be shown to the matcher, in order, and takes the first
14//! that a rule fires on. [`crate::term`] is what a way of showing one is, and the order is the
15//! most specific first: an operand that is a constant is offered as a constant before it is
16//! offered as a register, and an operand computed by an instruction of its own is offered as
17//! that instruction before it is offered as a register. A rule that wants an immediate too wide
18//! for the machine has a guard that turns it down, and the search carries on to the way of
19//! showing it that puts the constant in a register, which is the right answer and is one nobody
20//! had to write down.
21//!
22//! A constant is not lowered where it is written. It is materialized where a register for it is
23//! first wanted, which is what keeps a constant that every use folded into an immediate from
24//! leaving a dead instruction behind, and it also gives the value the shortest live range it
25//! could have. The instruction that materializes it comes from the rule set like everything else.
26//!
27//! # What it does not do yet
28//!
29//! Everything is in the general purpose registers, because every rule in the set is about an
30//! integer, so a call that passes a `double` and a function that returns one are both reported
31//! rather than lowered. So is an argument that travels on the stack, on either side of a call,
32//! and so is a call through an address rather than to a name.
33//!
34//! # A call
35//!
36//! Not a rule, because a rule pattern sees one term and what a call's operands are is whatever
37//! the signature made them. [`crate::abi`] builds one instead, out of the same description of the
38//! convention the arguments come from: the values it passes are reads constrained to the
39//! registers the convention places them in, what comes back is a write constrained to the
40//! register it comes back in, and every other register the callee is free to destroy is a write
41//! of that register and nothing else, which is all the allocator needs to keep a value out of it.
42//!
43//! What that costs the frame is an argument area, and nothing after selection could work out how
44//! big, so the size of the widest call is given back with the function. A function that makes no
45//! call at all is a leaf, and a leaf is the function that may use the red zone.
46//!
47//! # Where a block goes
48//!
49//! On the block, which is what machine IR does with an edge and is why the branches need no more
50//! rule language than the arithmetic did. A rule never names a block, so an unconditional jump
51//! has no rule at all and a conditional branch has one that is about its condition and nothing
52//! else. The arms are copied across after the block is filled, arguments and all, because an
53//! argument that is a constant is materialized where a register for it is first wanted and the
54//! end of the block is where an edge wants it.
55//!
56//! What this leaves behind is a function whose blocks are in the order the IR held them and whose
57//! branches are still branches on a register. Turning one into a `test` and a `jcc` is the block
58//! layout's, since which of the two arms falls through is the layout's answer, and [`crate::split`]
59//! has to run before allocation so that every edge carrying a value has somewhere to put it.
60//!
61//! A store and a return are the two things here that write no register. A store is emitted like
62//! everything else and the only difference is that there is no result to put anywhere, so the
63//! operands the target describes are all reads. A return is the same, and what it is for is its
64//! one operand: the target constrains it to the register the caller reads the value out of, and
65//! the allocator is what gets it there. The instruction that leaves is not chosen here at all,
66//! because the epilogue has to give the frame back first and [`crate::finish`] writes that after
67//! allocation, so a return of nothing is lowered to nothing.
68//!
69//! The entry block is the one block whose parameters are not block parameters here. They are the
70//! function's arguments, they are already somewhere when it starts, and [`crate::abi`] is what
71//! says where. An argument that arrives on the stack is reported rather than read, because where
72//! the stack put it is a distance into a frame and no frame exists until after allocation.
73//!
74//! Blocks are walked in the order the function holds them and a value is expected to be defined
75//! before it is used, which is true of the IR this is given because every pass before it keeps
76//! definitions ahead of uses.
77
78use std::collections::HashSet;
79use std::fmt;
80
81use rucc_base::{Interner, Symbol};
82use rucc_diag::Span;
83use rucc_ir::{
84 Abi, AsmOperands, Block, Def, Extra, FloatPred, Func, Inst, Linkage, MemOrder, Opcode, Param,
85 PrefetchHint, RmwOp, Type, Value, Visibility,
86};
87use rucc_mir as mir;
88use rucc_target::x86_64;
89use rucc_target::{CallRegs, Constraint, RegClass, Segment};
90
91use crate::abi::{self, Missing, Refused};
92use crate::coverage::Fired;
93use crate::elsewhere::Elsewhere;
94use crate::frame::{Layout, Local};
95use crate::select::{Match, Piece, Rule, Table};
96use crate::term::{MAX_ARGS, PLAIN, Plan, Shown, Term, Terms};
97use crate::varargs;
98
99/// The prefix a rule file puts in front of a machine term, which says which target it belongs
100/// to and is not part of the opcode.
101pub(crate) const PREFIX: &str = "x64.";
102
103/// The instruction a global offset table slot is read with.
104///
105/// Not in [`x86_64::FRAME`] with the other opcodes this file names, because a frame has no use for
106/// it. It is spelled out here because the relocation it takes is only legal on a `mov` with a REX
107/// prefix, so the width is part of the requirement rather than a choice.
108const GOT_LOAD: &str = "mov_rm_64";
109
110/// How wide an address is on this target, which is the width a cast between a pointer and an
111/// integer has to be at for the cast to be nothing.
112const ADDRESS_BITS: u32 = 64;
113
114/// How many bytes a `long double` takes in memory, and what it is aligned to, which are the same
115/// number and are both more than the ten bytes that mean anything.
116///
117/// The psABI's answer rather than a choice here. `sizeof (long double)` is sixteen on this
118/// machine, so an array of them is laid out this way whatever a slot holding one does, and a slot
119/// that agreed with the array is one fewer thing to get wrong.
120const X87_BYTES: u32 = 16;
121
122/// How many values the x87 stack holds at once.
123///
124/// Eight, which is the machine's number rather than a choice here, and it matters in one place:
125/// the parameters of a block are copied through the stack so that they all move at once, and a
126/// block with more of them than this has nowhere to put the ninth.
127const X87_DEPTH: usize = 8;
128
129/// How many bytes a value passes through on its way between a register and the x87 stack.
130///
131/// Eight, because the widest thing that crosses is a `double` or a sixty four bit integer, and
132/// nothing crosses at eighty bits: a value that wide is already in the frame and the stack reaches
133/// it where it is.
134const X87_CROSSING: u32 = 8;
135
136/// Where the rounding field of the x87 control word is and what it has to be set to for the unit
137/// to cut towards zero, which is the one rounding C asks for that the unit does not do by default.
138///
139/// Both bits on is truncate. The field is ORed into the word that was already there rather than
140/// written over it, so the precision control and the exception masks somebody else set stay set.
141const X87_TRUNCATE: i64 = 0x0c00;
142
143/// Whether a type is the one this machine has no register for.
144///
145/// Only the eighty bit float is, and that is a fact about x86-64 rather than about floats: every
146/// other scalar the front end produces is in a general purpose register or a vector one, and this
147/// one is on the x87 stack while it is being worked on and in memory the rest of the time. So it
148/// has no place in [`Lowering::class_of`] and no name in [`crate::term`], and every instruction
149/// that touches one is written out by hand in this file.
150fn on_x87(ty: Type) -> bool {
151 ty.is_scalar() && ty.is_float() && ty.bits() == 80
152}
153
154/// Why a function could not be lowered.
155///
156/// One reason and then nothing. A function with no rule for something in it is a function this
157/// cannot finish, and the second thing it could not lower is not news.
158#[derive(Debug, Clone, PartialEq, Eq)]
159pub enum Unsupported {
160 /// An instruction no rule fires on.
161 Inst {
162 /// The instruction that stopped it.
163 inst: Inst,
164 /// What the rule file would call it, or nothing if the rule language has no name for it
165 /// at all, which is what an instruction at a width nothing is written about looks like.
166 term: Option<&'static str>,
167 /// The opcode, which is what gets named when the rule language has no word for it.
168 ///
169 /// An opcode the rule language has no word for is exactly the opcode no rule lowers, so
170 /// without this the message would be empty in every case where somebody needs it.
171 opcode: Opcode,
172 /// What it produces, or nothing for an instruction that is only an effect.
173 ty: Option<Type>,
174 },
175 /// A parameter that does not arrive somewhere this can bring it in from.
176 ///
177 /// Not an instruction, which is why it is a separate arm: it is a fact about the signature
178 /// and there is nothing in the body of the function to point at.
179 Argument {
180 /// Its position in the signature.
181 index: usize,
182 /// What is wrong with where it arrives.
183 missing: Missing,
184 },
185 /// A call that passes or gives back a value this cannot put where the convention wants it.
186 Call {
187 /// The call.
188 inst: Inst,
189 /// Which value, and what is wrong with where it travels.
190 refused: Refused,
191 },
192 /// A `return` this cannot put where the convention wants it.
193 ///
194 /// A separate arm from [`Unsupported::Inst`] because it is not an instruction no rule fires
195 /// on. A return of more than one value is built from the convention rather than matched, the
196 /// same way a call is, so what goes wrong with one is what goes wrong with a call and not the
197 /// absence of a rule.
198 Returned {
199 /// The `return`.
200 inst: Inst,
201 /// What is wrong with where one of the values travels.
202 missing: Missing,
203 },
204 /// A stack slot the frame cannot give the bytes it asked for.
205 ///
206 /// Not an instruction no rule covers. An `alloca` is built here rather than matched, so what
207 /// goes wrong with one is what the frame can and cannot hold rather than what the rules spell.
208 Dynamic {
209 /// The `alloca`.
210 inst: Inst,
211 /// What the frame could not do about it.
212 growing: Growing,
213 },
214 /// More parameters of a type that travels on the x87 stack than the stack is deep.
215 ///
216 /// Not an instruction either, for the reason a function's parameter is not one: it is a fact
217 /// about the block and there is nothing in the block to point at. What crosses an edge for one
218 /// of these is the address of where the value is, and the block copies the bytes into a slot
219 /// of its own, all of them through the stack at once so that a block carrying two of them
220 /// swapped is copied in an order that is right. Eight is as many as the stack holds, and a
221 /// ninth would have to be copied before or after the rest, which is the order that could be
222 /// wrong.
223 Phi {
224 /// Which block it arrives at.
225 block: Block,
226 /// How many of them arrive there, which is the whole of what is wrong.
227 count: usize,
228 /// What they are.
229 ty: Type,
230 },
231 /// An `asm` statement this cannot build.
232 ///
233 /// Not an instruction no rule fires on, for the reason a call is not one: what it stands for is
234 /// whatever its template says, and no pattern over terms can read a string.
235 Assembly {
236 /// The `inline_asm`.
237 inst: Inst,
238 /// What about it is not built here yet.
239 refused: Written,
240 },
241}
242
243/// What about an `asm` statement is not built yet.
244#[derive(Debug, Clone, Copy, PartialEq, Eq)]
245pub enum Written {
246 /// A template with instructions in it.
247 Template,
248 /// An `asm goto`, whose labels make the statement a terminator.
249 Goto,
250 /// An operand this cannot put where the constraint says it goes.
251 Operand,
252}
253
254impl Written {
255 /// The rest of the sentence that starts with the statement.
256 #[must_use]
257 pub fn why(self) -> &'static str {
258 match self {
259 // The template is the assembler's to read and there is no assembler here yet, so a
260 // template with anything in it is a string nothing can turn into bytes. An empty one is
261 // no instructions, and no instructions is something this can write.
262 Written::Template => "has instructions in its template, which nothing here assembles",
263 Written::Goto => "jumps to a label, which nothing here builds an edge for",
264 Written::Operand => "has an operand this cannot place",
265 }
266 }
267}
268
269/// What the frame could not do about a stack slot.
270#[derive(Debug, Clone, Copy, PartialEq, Eq)]
271pub enum Growing {
272 /// An object of a size the number a frame counts bytes in does not reach.
273 Huge,
274 /// A variable length array wanting more alignment than a call leaves the stack pointer with.
275 ///
276 /// Rounding the stack pointer down again after the bytes have been taken would put it
277 /// somewhere no constant reaches the rest of the frame from, so a frame like this needs a
278 /// second base register held for the whole of the function. Nothing here holds one.
279 Aligned,
280 /// A variable length array in a function whose frame is meant to be touched a page at a time.
281 ///
282 /// The pages a prologue takes are touched by the prologue, which knows how many there are when
283 /// it is written. The pages a variable length array takes are not known until the declaration
284 /// runs, so touching them is a loop next to the declaration, and there is no loop here yet.
285 Probed,
286}
287
288impl Growing {
289 /// The rest of the sentence that starts with the slot.
290 #[must_use]
291 pub fn why(self) -> &'static str {
292 match self {
293 Growing::Huge => "is more bytes than a frame counts",
294 Growing::Aligned => {
295 "wants more alignment than the stack pointer is left on, which needs a base \
296 register nothing here keeps"
297 }
298 Growing::Probed => {
299 "grows the stack, and nothing here touches the pages it takes a page at a time"
300 }
301 }
302 }
303}
304
305impl Unsupported {
306 /// The instruction it is about, or nothing for the one arm that is about a signature.
307 ///
308 /// What a caller wants this for is the span. The function knows where every instruction in
309 /// it came from, so a caller holding both can point a message at the line somebody wrote
310 /// rather than at the file as a whole, and nothing here has to carry a span of its own.
311 pub fn inst(&self) -> Option<Inst> {
312 match *self {
313 Unsupported::Inst { inst, .. }
314 | Unsupported::Call { inst, .. }
315 | Unsupported::Returned { inst, .. }
316 | Unsupported::Dynamic { inst, .. }
317 | Unsupported::Assembly { inst, .. } => Some(inst),
318 Unsupported::Argument { .. } | Unsupported::Phi { .. } => None,
319 }
320 }
321}
322
323impl fmt::Display for Unsupported {
324 fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
325 match *self {
326 Unsupported::Inst { term: Some(term), .. } => write!(f, "no rule lowers `{term}`"),
327 Unsupported::Inst { term: None, opcode, ty: Some(ty), .. } => {
328 write!(f, "no rule lowers a `{opcode}` producing a `{ty}`")
329 }
330 Unsupported::Inst { term: None, opcode, ty: None, .. } => {
331 write!(f, "no rule lowers a `{opcode}`")
332 }
333 Unsupported::Argument { index, missing } => {
334 write!(f, "parameter {index} {}", missing.why())
335 }
336 Unsupported::Call { refused: Refused { argument: Some(index), missing }, .. } => {
337 write!(f, "argument {index} of this call {}", missing.why())
338 }
339 Unsupported::Call { refused: Refused { argument: None, missing }, .. } => {
340 write!(f, "what this call gives back {}", missing.why())
341 }
342 Unsupported::Returned { missing, .. } => {
343 write!(f, "what this function gives back {}", missing.why())
344 }
345 Unsupported::Dynamic { growing, .. } => {
346 write!(f, "this local {}", growing.why())
347 }
348 Unsupported::Phi { block, count, ty } => {
349 let block = block.index();
350 write!(
351 f,
352 "block{block} takes {count} parameters of type `{ty}` and only {X87_DEPTH} can cross an edge at once"
353 )
354 }
355 Unsupported::Assembly { refused, .. } => write!(f, "this `asm` {}", refused.why()),
356 }
357 }
358}
359
360impl std::error::Error for Unsupported {}
361
362/// A lowered function, and what the frame needs that the machine IR does not hold.
363#[derive(Debug)]
364pub struct Lowered {
365 /// The function, in machine instructions.
366 pub func: mir::Func,
367 /// What it wants its stack to look like, which is separate from the function so that the two
368 /// can be read and written at the same time.
369 pub stack: Stack,
370 /// Which rules of the table lowered it, which is what `-Zrule-coverage` asks for and what
371 /// `crate::coverage` writes down.
372 pub fired: Fired,
373 /// Which machine IR block each IR block became, indexed by the IR block's own index, and
374 /// nothing for a block the walk never reached.
375 ///
376 /// Here because it is the only place the correspondence exists. Selection makes one block per
377 /// block, in the same order and with the arms in the same order, so anything the IR knows
378 /// about a block can be carried down through this and nothing else, and
379 /// [`crate::weights::carry`] is what does.
380 pub blocks: Vec<Option<mir::Block>>,
381}
382
383/// What a function's stack has to hold, as far as selection is able to say.
384///
385/// All of it is answered here because selection is where a call is built and where an `alloca`
386/// is read, and nothing after it could tell what either of them needed.
387#[derive(Debug, Default)]
388pub struct Stack {
389 /// How many bytes the widest call in the function needs below the stack pointer for the
390 /// arguments it passes there, or `None` for a function that makes no call at all.
391 ///
392 /// `None` is a leaf, which is the function that may use the red zone and the one whose stack
393 /// pointer does not have to be left aligned for anybody.
394 pub calls: Option<u32>,
395 /// The memory the function asked for itself, one entry for every `alloca` in it, in the order
396 /// the walk reached them.
397 pub locals: Vec<Local>,
398 /// Which instruction computes the address of which of those locals.
399 ///
400 /// An address in the frame is a distance from the stack pointer, and there is no frame until
401 /// after allocation, so the instruction is written here with nothing in its displacement and
402 /// [`crate::finish`] writes the number in once [`crate::frame::Frame`] knows it.
403 pub addresses: Vec<(mir::Inst, usize)>,
404 /// Which instruction computes the address of a piece of memory whose size the function works
405 /// out while it runs, which is what a variable length array is.
406 ///
407 /// Waiting on [`crate::finish`] for a different number from the one the addresses above are:
408 /// the bytes were taken off the stack pointer by the instruction in front of this one, so where
409 /// they start is however much of the bottom of the frame belongs to the arguments of a call,
410 /// and that is not known until the frame is.
411 pub dynamic: Vec<mir::Inst>,
412 /// Where the function first moves the stack pointer while it runs, if it does at all.
413 ///
414 /// Two things are read off this. One is whether at all, which is what [`crate::frame::Layout`]
415 /// wants, because a frame that moves its stack pointer has a different shape from one that does
416 /// not and the layout is built before the instructions are looked at again. See `Growing` in
417 /// [`crate::frame`]. The other is where, so that a caller that cannot accept such a frame has
418 /// somewhere to point when it says so.
419 pub grown_at: Option<Inst>,
420 /// Which instruction reads which of the arguments the caller passed on the stack, as how far up
421 /// the caller's argument area it reads.
422 ///
423 /// Waiting on [`crate::finish`] for the same reason the addresses above are, and on one thing
424 /// more: where the caller's argument area is from inside this function depends on whether the
425 /// prologue had to force the stack pointer's alignment, so which register the load reads
426 /// through is not settled here either.
427 pub arguments: Vec<(mir::Inst, u32)>,
428}
429
430impl Stack {
431 /// The layout given, with the three fields only the lowering knows the answer to filled in.
432 ///
433 /// Everything else in a layout comes from the flags the function is compiled under or from the
434 /// allocation, so this takes one and returns it rather than building one.
435 #[must_use]
436 pub fn layout<'a>(&'a self, base: Layout<'a>) -> Layout<'a> {
437 Layout {
438 leaf: self.calls.is_none(),
439 outgoing: self.calls.unwrap_or(0),
440 locals: &self.locals,
441 grows: self.grown_at.is_some(),
442 ..base
443 }
444 }
445}
446
447/// The x86-64 machine IR for that function.
448///
449/// # Errors
450///
451/// The first instruction no rule fires on, which today is anything at a width the rule set is not
452/// written at, a parameter that does not arrive in a register this can read, or a call that
453/// passes something this cannot put where the convention wants it.
454pub fn func(
455 source: &Func,
456 names: &mut Interner,
457 conv: &'static CallRegs,
458 elsewhere: &Elsewhere,
459) -> Result<Lowered, Unsupported> {
460 Lowering::new(source, names, conv, elsewhere).run()
461}
462
463/// What the matcher settled on for one block, indexed the way the block's instructions are.
464struct Decided {
465 /// What each instruction matched, and nothing for one that matched no rule or was folded
466 /// into a later one.
467 found: Vec<Option<Match<Term>>>,
468 /// How each instruction showed its operands to the matcher, which is what says what it took.
469 plans: Vec<Option<Plan>>,
470 /// The instructions some other instruction took, which are the ones with nothing to write.
471 folded: Vec<Inst>,
472}
473
474/// One function being lowered.
475struct Lowering<'a> {
476 source: &'a Func,
477 names: &'a mut Interner,
478 out: mir::Func,
479 /// The machine register each IR value is in, once it has one.
480 regs: Vec<Option<mir::Reg>>,
481 /// For a constant that has been written into a register, the block it was written into,
482 /// which is the only block that register is any good in.
483 written: Vec<Option<mir::Block>>,
484 /// How many times each IR value is read, which is what says whether an instruction may be
485 /// folded into the one that reads it.
486 uses: Vec<u32>,
487 /// The block being filled.
488 at: Option<mir::Block>,
489 /// The machine IR block each IR block became.
490 blocks: Vec<Option<mir::Block>>,
491 /// The class an address is in, which is the general purpose one and is not a question: every
492 /// register an addressing mode names holds part of an address, and there is no machine here
493 /// that computes an address anywhere but in this file. Which class a *value* is in is
494 /// [`Lowering::class_of`], and it is a question, because a float is in the other one.
495 gpr: RegClass,
496 /// Where the convention this function is compiled for puts things, which is read for the
497 /// arguments and for the calls.
498 conv: &'static CallRegs,
499 /// Which names this function may not work an address out for itself, which is a fact about the
500 /// module and so is worked out before any of this and handed in.
501 elsewhere: &'a Elsewhere,
502 /// What the function wants its stack to look like, filled in as the walk finds out.
503 stack: Stack,
504 /// What a `va_start` in this function has to write, or nothing for a function that takes no
505 /// arguments its signature does not name.
506 ///
507 /// Worked out once, when the entry block binds the parameters, because every number in it is
508 /// about where those parameters left the walk over the argument registers and there is nowhere
509 /// else that knows.
510 varargs: Option<Varargs>,
511 /// Which of the function's stack objects each eighty bit value lives in, once it has asked
512 /// for one.
513 ///
514 /// One slot per value and it is never given back, which is what makes an eighty bit value
515 /// behave like every other one: it is written once and read wherever it is read, and no two
516 /// of them share a slot the way two of them would share a register. What is in a register is
517 /// the address, and that is worked out again at every use rather than kept, so nothing here
518 /// holds a general purpose register open across a whole function.
519 slots: Vec<Option<usize>>,
520 /// The eight bytes a value passes through between a register and the x87 stack, once
521 /// something has wanted them.
522 ///
523 /// One for the whole function, because every group that uses it is a handful of instructions
524 /// with nothing in between: the bytes are written, read straight back and never looked at
525 /// again, so a second slot would be a second slot holding the same nothing.
526 crossing: Option<usize>,
527 /// The four bytes the control word is saved in and the changed copy written to, once
528 /// something has wanted them.
529 ///
530 /// One for the whole function for the reason above, and four rather than two because it is
531 /// two words: the one the unit had and the one with the rounding field turned to truncate.
532 control: Option<usize>,
533 /// Which rules have fired so far.
534 fired: Fired,
535}
536
537/// What a `va_start` in a variadic function writes into the list it is given.
538///
539/// Three of the four are settled here and the fourth is not a number at all yet: where the save
540/// area is and where the caller's argument area is are both distances into a frame that does not
541/// exist until after allocation, so both are `lea` instructions [`crate::finish`] fills in.
542#[derive(Debug, Clone, Copy, PartialEq, Eq)]
543struct Varargs {
544 /// Which of the function's stack objects is the register save area.
545 save: usize,
546 /// How far up the caller's argument area the first argument the signature does not name is,
547 /// which is the whole of that area the named ones did not take.
548 incoming: u32,
549 /// What `gp_offset` starts at, which is past the general purpose registers the named arguments
550 /// took.
551 integers: u32,
552 /// What `fp_offset` starts at, which is past the vector ones.
553 floats: u32,
554}
555
556/// How far a function's name reaches, narrowed from the linkage the IR gave it.
557///
558/// The IR has five and an object file says three, and the two the linker cannot tell apart are
559/// the two weak ones: which of them a symbol had is a fact the optimizer reads and the linker has
560/// no way to record. A function is never `Common`, since that is what a tentative definition of an
561/// object is and there is no tentative definition of a function, and it is written here rather
562/// than left out so that a linkage added later has to come past this.
563const fn binding(linkage: Linkage) -> mir::Binding {
564 match linkage {
565 Linkage::Internal => mir::Binding::Local,
566 Linkage::Weak | Linkage::LinkOnce => mir::Binding::Weak,
567 Linkage::External | Linkage::Common => mir::Binding::Global,
568 }
569}
570
571/// How far a function's name reaches outside a shared library, carried across unchanged.
572///
573/// Nothing is narrowed here the way [`binding`] narrows the linkage, because ELF records all
574/// three of these and the two enumerations are the same three answers written twice: once in a
575/// crate that is not allowed to know what an object file is and once in one that is.
576const fn visibility(visibility: Visibility) -> mir::Visibility {
577 match visibility {
578 Visibility::Default => mir::Visibility::Default,
579 Visibility::Hidden => mir::Visibility::Hidden,
580 Visibility::Protected => mir::Visibility::Protected,
581 }
582}
583
584impl<'a> Lowering<'a> {
585 fn new(
586 source: &'a Func,
587 names: &'a mut Interner,
588 conv: &'static CallRegs,
589 elsewhere: &'a Elsewhere,
590 ) -> Self {
591 let counts = source.counts();
592 let name = source.name;
593 let mut uses = vec![0; counts.values];
594 for block in source.blocks() {
595 for inst in source.insts(block) {
596 for &arg in &source[source[inst].args] {
597 uses[arg.index()] += 1;
598 }
599 for call in source.successors(inst) {
600 for &arg in &source[call.args] {
601 uses[arg.index()] += 1;
602 }
603 }
604 }
605 }
606 let mut out = mir::Func::new(name);
607 out.align = source.align;
608 out.binding = binding(source.linkage);
609 out.visibility = visibility(source.visibility);
610 Self {
611 source,
612 names,
613 out,
614 regs: vec![None; counts.values],
615 written: vec![None; counts.values],
616 blocks: vec![None; counts.blocks],
617 uses,
618 at: None,
619 gpr: x86_64::GPR,
620 conv,
621 elsewhere,
622 stack: Stack::default(),
623 varargs: None,
624 slots: vec![None; counts.values],
625 crossing: None,
626 control: None,
627 fired: Fired::new(),
628 }
629 }
630
631 fn run(mut self) -> Result<Lowered, Unsupported> {
632 // Every block before any of them is filled, because a block that jumps forward has to
633 // name the block it jumps to and a machine IR block is named by a handle rather than by
634 // the IR block it came from.
635 for block in self.source.blocks() {
636 let out = self.out.create_block();
637 self.blocks[block.index()] = Some(out);
638 }
639 for block in self.order() {
640 self.block(block)?;
641 }
642 Ok(Lowered { func: self.out, stack: self.stack, fired: self.fired, blocks: self.blocks })
643 }
644
645 /// The order the blocks are filled in, which is not the order they are written in.
646 ///
647 /// Reverse postorder, because a value is written in a block that dominates every block that
648 /// reads it and a block in reverse postorder comes before every block it dominates. The order
649 /// the blocks are written in does not have that property: a block written early can read a
650 /// value a block below it writes, and reading a value with no register yet mints one, so the
651 /// register the definition writes later is not the register the read named. Nothing writes the
652 /// one the read named, and what comes out is a function that loads a stack slot no store ever
653 /// reached. It is the order this walk goes in rather than the order the blocks come out in,
654 /// which is what the loop above fixes, so the machine function is still written the way the IR
655 /// function was.
656 ///
657 /// Blocks the entry does not reach come last, in the order they are written in. Nothing runs
658 /// them and nothing they name is read by anything that does, but they still have to be filled,
659 /// because a machine block with no terminator is not one the passes below can read.
660 fn order(&self) -> Vec<Block> {
661 let Some(entry) = self.source.entry() else { return self.source.blocks().collect() };
662 let count = self.blocks.len();
663 let mut succs: Vec<Vec<Block>> = vec![Vec::new(); count];
664 for block in self.source.blocks() {
665 let Some(term) = self.source.terminator(block) else { continue };
666 succs[block.index()] = self.source.successors(term).map(|call| call.block).collect();
667 }
668 // An explicit stack, because the depth of the walk is the number of blocks and a function
669 // built by a generator has as many of those as it likes.
670 let mut seen = vec![false; count];
671 let mut order = Vec::with_capacity(count);
672 let mut stack = vec![(entry, 0usize)];
673 seen[entry.index()] = true;
674 while let Some((block, at)) = stack.pop() {
675 let Some(&next) = succs[block.index()].get(at) else {
676 order.push(block);
677 continue;
678 };
679 stack.push((block, at + 1));
680 if !seen[next.index()] {
681 seen[next.index()] = true;
682 stack.push((next, 0));
683 }
684 }
685 order.reverse();
686 order.extend(self.source.blocks().filter(|block| !seen[block.index()]));
687 order
688 }
689
690 /// One block: its parameters, then every instruction in it that is not folded into another.
691 fn block(&mut self, block: Block) -> Result<(), Unsupported> {
692 let out = self.out_block(block);
693 self.at = Some(out);
694 if self.source.entry() == Some(block) {
695 self.arrive(block, out)?;
696 } else {
697 let mut arriving = Vec::new();
698 for ¶m in &self.source[block].params {
699 // A value with no register to arrive in, which the class would not say, since
700 // `class_of` puts one of these in the general purpose file on purpose and what it
701 // means by that is that nothing there can hold it. What crosses the edge for one
702 // of those is the address of where the value already is, so the parameter is a
703 // pointer here and the bytes it points at are copied below.
704 let ty = self.source[param].ty;
705 let reg = self.out.append_param(out, self.class_of(ty));
706 self.regs[param.index()] = Some(reg);
707 if on_x87(ty) {
708 arriving.push((param, reg));
709 }
710 }
711 self.settle(block, &arriving)?;
712 }
713
714 // What each instruction matched, and which instructions were folded into another. The
715 // decision is made for the whole block before any of it is written, and it is made more
716 // than once: a value that only some of its readers took has to be put back in a register
717 // for all of them, and taking it away from those readers changes what they match.
718 let insts: Vec<Inst> = self.source.insts(block).collect();
719 let mut refused: HashSet<Value> = HashSet::new();
720 let mut decided = self.decide(&insts, &refused);
721 while let Some(value) = self.left_alive(&insts, &decided.plans) {
722 refused.insert(value);
723 decided = self.decide(&insts, &refused);
724 }
725 let Decided { found, folded, .. } = decided;
726
727 for (&inst, matched) in insts.iter().zip(found) {
728 if folded.contains(&inst) || self.writes_nothing(inst) {
729 continue;
730 }
731 // A call is built from the convention rather than matched, which is why it is the one
732 // opcode looked at by name here. Through an address it is a different instruction and
733 // the same convention, so the two arrive at the same place and differ in one line of
734 // it.
735 match self.source[inst].opcode {
736 Opcode::Call | Opcode::CallIndirect => {
737 self.called(inst)?;
738 continue;
739 }
740 // Built from the frame rather than matched, for the same shape of reason a call
741 // is built from the convention: what a rule replaces a term with is instructions,
742 // and what an `alloca` needs first is bytes, which the rule language has no way
743 // to ask for.
744 Opcode::Alloca => {
745 self.reserve(inst)?;
746 continue;
747 }
748 // Reading the stack pointer and writing it back, which are the two ends of a scope
749 // holding a variable length array. Built here for the reason an `alloca` is: the
750 // value is a register the rule language has no way to name, because what it holds
751 // is not a value the program computed but where the machine's stack had got to.
752 Opcode::StackSave => {
753 self.stack_pointer(inst, false)?;
754 continue;
755 }
756 Opcode::StackRestore => {
757 self.stack_pointer(inst, true)?;
758 continue;
759 }
760 // The address of a name, built here for the same reason an `alloca` is: what a
761 // rule replaces a term with is instructions over values, and the operand of this
762 // one is a symbol, which is a thing the rule language has no way to bind and the
763 // solver has no way to say anything about. There is nothing in `lea sym(%rip)` a
764 // proof over bitvectors could discharge, because what makes it the right answer
765 // is the relocation and what the linker does with it.
766 Opcode::GlobalAddr => {
767 self.address_of(inst)?;
768 continue;
769 }
770 // Where this thread's own storage starts, built here for a reason of the same
771 // shape: what it reads is `%fs`, which is not a register the rule language can
772 // bind and not one a proof over bitvectors could say anything about, because what
773 // makes the load the right answer is an agreement between the loader and the C
774 // library rather than any arithmetic.
775 Opcode::ThreadPointer => {
776 self.thread_pointer(inst)?;
777 continue;
778 }
779 // Built from the frame for the reason an `alloca` is, and from the convention for
780 // the reason a call is: three of the four fields it writes are distances that do
781 // not exist until the frame does, and the fourth is where the walk over the
782 // argument registers stopped. A function that is not variadic has no such walk to
783 // report, so it has nothing here and is refused below, which is the right answer
784 // for a `va_start` in one.
785 Opcode::VaStart if self.varargs.is_some() => {
786 self.va_start(inst)?;
787 continue;
788 }
789 // A return of more than one value, which is a structure small enough to come
790 // back in a pair of registers. Built from the convention for the reason a call
791 // is: which register each half goes in depends on the halves in front of it,
792 // because the two register files are walked separately, and a pattern over a term
793 // cannot see them. A return of one value is a term with a name and a rule, and it
794 // stays one.
795 //
796 // A return of none in a function whose answer went through memory is here too,
797 // and for a different reason: what it gives back is not written in the IR at all.
798 // The convention says the address the caller handed over comes back, and only the
799 // signature says this function was handed one.
800 //
801 // And a return of one eighty bit value, for a third reason: what a rule would
802 // write is an instruction leaving the value in a register, and this one is left on
803 // the x87 stack instead. A rule could not name that stack any more than any other
804 // rule about this type could.
805 Opcode::Return
806 if self.source[self.source[inst].args].len() > 1
807 || self.sret().is_some()
808 || self.gives_back_x87(inst) =>
809 {
810 self.returned(inst)?;
811 continue;
812 }
813 // A cast between a pointer and an integer of the same width, which on this
814 // machine is every one the front end writes. No instruction at all, so no rule
815 // could name one.
816 Opcode::PtrToInt | Opcode::IntToPtr => {
817 self.rename(inst)?;
818 continue;
819 }
820 // A barrier, which is one instruction or none depending on the ordering. Written
821 // by name because there is nothing about it a rule could be proved against, the
822 // way there is nothing to prove about the address of a symbol.
823 Opcode::Fence => {
824 self.barrier(inst)?;
825 continue;
826 }
827 // A hint, written by name for the reason a barrier is and one step further: not
828 // only is there no equality for a proof to discharge, there is nothing about the
829 // program around it either. Which of the four instructions it is comes out of the
830 // number the builtin was given, which is beside the instruction rather than in it.
831 Opcode::Prefetch => {
832 self.hint(inst)?;
833 continue;
834 }
835 // A compare and exchange, which is written by name because it produces two values
836 // and a rule produces one. The replacement of a rule is one term, a term names the
837 // value an instruction computes, and there is no way in that language to say that
838 // an instruction leaves an answer in one place and a yes or no in another.
839 Opcode::Cmpxchg => {
840 self.exchange(inst)?;
841 continue;
842 }
843 // A read modify write, which is written by name for a different reason: it produces
844 // one value, so a rule could name it, and what it does is not in the head a rule
845 // matches on. Every one of the thirteen operations is the same opcode at the same
846 // type and differs only in what is carried beside it, so one pattern would be all
847 // thirteen patterns. Of the thirteen only the three with an instruction reach here,
848 // since `crate::retry` turned the rest into loops a long way above this.
849 Opcode::AtomicRmw => {
850 self.modify(inst)?;
851 continue;
852 }
853 // An `asm` statement, whose lowering is its template and there is no term for a
854 // string. Written by name for the reason a barrier is, and before the x87 arm
855 // below so that an `asm` holding a `long double` is refused as the `asm` it is
856 // rather than as an instruction nothing computes.
857 Opcode::InlineAsm => {
858 self.assembly(inst)?;
859 continue;
860 }
861 // Anything at all with an eighty bit float in it, which is the one arm here
862 // chosen by a type rather than by an opcode, because what makes these different
863 // is not what they do but where the value is. A `long double` has no register,
864 // so it has no name in `crate::term` and no rule could bind one: every one of
865 // these is a group of instructions over a frame slot, written out below.
866 //
867 // Last of the arms, so that a call and a return with one of these in them reach
868 // the convention first and are refused by it, which is the truer answer: what is
869 // wrong there is where the value has to travel and not that nothing can compute
870 // it.
871 _ if self.touches_x87(inst) => {
872 self.x87(inst)?;
873 continue;
874 }
875 _ => {}
876 }
877 let matched = matched.ok_or_else(|| self.unsupported(inst))?;
878 self.emit(inst, &matched)?;
879 // After it is built rather than when it matched, so that what is recorded is the rules
880 // this function was lowered by and not the rules something was tried with.
881 self.fired.mark(matched.rule);
882 }
883 self.edges(block, out)
884 }
885
886 /// One call, which is built from the convention rather than matched against the table for the
887 /// same reason the arguments of the function itself are.
888 ///
889 /// The arguments are read before the call is built, which is what materializes a constant
890 /// argument into a register, since no call passes an immediate.
891 ///
892 /// A call to a name and a call through an address are both here, and what tells them apart is
893 /// the opcode rather than whether a callee was recorded, which is the same thing the verifier
894 /// reads. Through an address the first operand is the address and the arguments are the ones
895 /// behind it, and everything after that is the same: where each argument goes, where the value
896 /// comes back and which registers are gone across it are the convention's answers and the
897 /// convention does not ask what is being called.
898 fn called(&mut self, inst: Inst) -> Result<(), Unsupported> {
899 let data = &self.source[inst];
900 let Extra::Call(info) = data.extra else { return Err(self.unsupported(inst)) };
901 let info = self.source[info];
902 let indirect = data.opcode == Opcode::CallIndirect;
903
904 let values: Vec<Value> = self.source[data.args].to_vec();
905 let callee = if indirect {
906 let &address = values.first().ok_or_else(|| self.unsupported(inst))?;
907 abi::Callee::Through(self.reg_of(address)?)
908 } else {
909 abi::Callee::Named(info.callee.ok_or_else(|| self.unsupported(inst))?)
910 };
911
912 // What the ABI asks of each argument, read out before any of them is, because reading one
913 // borrows the function this is a table in. The ones the signature names are the signature's
914 // answer and the ones behind them are the call's, which is where a structure passed to a
915 // variadic callee by value says that its bytes travel: there is no parameter to say it on.
916 let signature = &self.source[info.signature];
917 let variadic = signature.variadic;
918 let named: Vec<Abi> = signature.params.iter().map(|param| param.abi).collect();
919 let beyond: Vec<Abi> = self.source[info.varargs].to_vec();
920 // Every value that comes back and not only the first. A structure small enough to travel
921 // in registers comes back in up to two of them, and which register each half is in is the
922 // convention's answer, which is why the whole list goes to the same place the arguments do
923 // rather than to a rule.
924 let returns: Vec<Type> = signature.return_types().collect();
925
926 let mut args = Vec::with_capacity(values.len());
927 for (index, value) in values.into_iter().skip(usize::from(indirect)).enumerate() {
928 let abi = named.get(index).or_else(|| beyond.get(index - named.len()));
929 let abi = abi.copied().unwrap_or_default();
930 let ty = self.source[value].ty;
931 // What travels for an eighty bit value is its bytes, so what the call is handed is
932 // where they are rather than a register they are in, and there is no register they
933 // could be in. Everything else about it is a sixteen byte object passed by value and
934 // is built by the same code.
935 let reg =
936 if abi::on_the_stack(ty) { self.x87_slot(value) } else { self.reg_of(value)? };
937 args.push(abi::Passing { ty, reg, abi });
938 }
939 let block = self.at.expect("a block is being filled");
940 let what = abi::Calling { callee, args: &args, returns: &returns, variadic };
941 let made = abi::call(&mut self.out, block, &what, self.conv, self.names)
942 .map_err(|refused| Unsupported::Call { inst, refused })?;
943 let calls = &mut self.stack.calls;
944 *calls = Some(calls.unwrap_or(0).max(made.outgoing));
945 // An eighty bit value came back on the x87 stack, and the one thing that has to happen
946 // before anything else touches that stack is taking it off. So the `fstp` goes here, in
947 // front of everything the block does next, and after it the value is in its slot and is
948 // read the way every other one is.
949 let results: Vec<Value> = self.source[inst].results().collect();
950 if let [result] = results[..] {
951 if abi::on_the_stack(self.source[result].ty) {
952 let span = self.source.span(inst);
953 let into = self.x87_slot(result);
954 let into = self.through(into);
955 self.x87_at("fstp_t", span, into);
956 return Ok(());
957 }
958 }
959 for (result, ®) in results.into_iter().zip(&made.results) {
960 self.regs[result.index()] = Some(reg);
961 }
962 Ok(())
963 }
964
965 /// The pointer a function returning through memory was handed, or nothing in a function that
966 /// was not.
967 ///
968 /// It is the first parameter and the signature is what says so, since in the IR it is an
969 /// ordinary pointer and reads like one everywhere in the body. A function with a signature
970 /// like that and no entry block has nothing to give back and no body to give it back from.
971 fn sret(&self) -> Option<Value> {
972 let first = self.source.signature().params.first()?;
973 if !matches!(first.abi, Abi::Sret { .. }) {
974 return None;
975 }
976 self.source[self.source.entry()?].params.first().copied()
977 }
978
979 /// One `return` the convention has to write, as the place each value has to be in by the end.
980 ///
981 /// One pseudo per value, each a read constrained to a return register, which is what a return
982 /// of one value already is and is the whole of what either does. The `ret` itself comes from
983 /// the epilogue for both, long after this, because the frame has to be given back first.
984 ///
985 /// The two register files are counted separately, so a structure of a `double` and a `long`
986 /// leaves the `double` in the first vector register and the `long` in the first integer one
987 /// rather than in the second of either. That is the same walk `rucc_codegen::abi` makes on
988 /// the other side of the call, which is what makes the two ends agree.
989 ///
990 /// A function whose answer went through memory gives back the address it was handed, in front
991 /// of nothing else, because a signature that returns that way returns nothing else. That the
992 /// caller already knows the address is not enough: it is allowed to read the register instead,
993 /// and a caller that does gets whatever the allocator last left there. In a leaf function that
994 /// is usually the right answer by accident, and one call in the body is enough to make it a
995 /// wild pointer, which is why this is written rather than left to luck.
996 ///
997 /// Where everything goes is worked out before anything is written, so a return this cannot
998 /// make leaves no half of one behind.
999 /// Whether what a `return` gives back is the one value that goes back on the x87 stack.
1000 fn gives_back_x87(&self, inst: Inst) -> bool {
1001 let [value] = self.source[self.source[inst].args] else { return false };
1002 abi::on_the_stack(self.source[value].ty)
1003 }
1004
1005 fn returned(&mut self, inst: Inst) -> Result<(), Unsupported> {
1006 let values: Vec<Value> = self.source[self.source[inst].args].to_vec();
1007 let (mut ints, mut floats) = (0usize, 0usize);
1008 let mut parts = Vec::with_capacity(values.len() + 1);
1009 // An eighty bit value goes back on the x87 stack, which is where the convention says it is
1010 // and is the one place a value is left rather than put in a register. So the whole of the
1011 // return is an `fld` of its slot, and the stack it leaves the value on is not empty at the
1012 // `ret`, which is the one time in this file that is true and is what the convention asks
1013 // for. What comes after is the epilogue, which gives the frame back and touches nothing in
1014 // the unit.
1015 if let [value] = values[..] {
1016 let ty = self.source[value].ty;
1017 if abi::on_the_stack(ty) && self.sret().is_none() {
1018 let span = self.source.span(inst);
1019 let from = self.x87_slot(value);
1020 let from = self.through(from);
1021 self.x87_at("fld_t", span, from);
1022 return Ok(());
1023 }
1024 }
1025 for value in self.sret().into_iter().chain(values) {
1026 let ty = self.source[value].ty;
1027 let at = if crate::term::float_slot(ty).is_some() { &mut floats } else { &mut ints };
1028 // Why it cannot come back, and not only that it cannot. A type that travels nowhere
1029 // says so itself, and a type that travels perfectly well ran out of registers.
1030 let missing = abi::refuses(ty).unwrap_or(Missing::NoRoom);
1031 let name = abi::ret_of(ty, *at).ok_or(Unsupported::Returned { inst, missing })?;
1032 *at += 1;
1033 // The register is the target's answer and not one worked out here, the same as it is
1034 // for a return of one value, so that both halves of a pair and every rule that writes
1035 // half of one are reading the same table.
1036 let opcode = name.strip_prefix(PREFIX).expect("a machine instruction of this target");
1037 let form = x86_64::form(opcode).ok_or_else(|| self.unsupported(inst))?;
1038 let [desc] = form.operands() else { return Err(self.unsupported(inst)) };
1039 parts.push((self.names.intern(name), self.reg_of(value)?, *desc));
1040 }
1041
1042 let block = self.at.expect("a block is being filled");
1043 let span = self.source.span(inst);
1044 for (opcode, reg, desc) in parts {
1045 let operand = mir::Operand {
1046 reg,
1047 class: desc.class,
1048 role: desc.role,
1049 constraint: desc.constraint,
1050 };
1051 self.out.build(block, mir::Opcode::new(opcode)).at(span).operand(operand).finish();
1052 }
1053 Ok(())
1054 }
1055
1056 /// One `alloca`: the bytes it asks for go on the list the frame is laid out from, and the
1057 /// address of them is one instruction.
1058 ///
1059 /// The instruction is a `lea` off the stack pointer, which is the one register that reaches
1060 /// the frame in every function, and its displacement is left at nothing because there is no
1061 /// frame yet. Which instruction is waiting for which local is remembered, and
1062 /// [`crate::finish`] fills the numbers in after [`crate::frame::Frame`] has placed them.
1063 ///
1064 /// There is deliberately no rule for `alloca` and no name for one in [`crate::term`], and
1065 /// that is what stops it being folded into something else. An operand shown as the
1066 /// instruction that computed it is offered to the matcher by its name, so an `alloca` with no
1067 /// name is one no pattern can reach past, and the address it computes is always in a register
1068 /// by the time anything reads it.
1069 fn reserve(&mut self, inst: Inst) -> Result<(), Unsupported> {
1070 let data = &self.source[inst];
1071 // A variable length array carries the size it wants as an operand rather than in the
1072 // instruction, which is the whole of what tells the two apart here.
1073 if let Some(&size) = self.source[data.args].first() {
1074 return self.grow(inst, size);
1075 }
1076 let Extra::Mem(mem) = data.extra else { return Err(self.unsupported(inst)) };
1077 let info = self.source[mem];
1078 let size = u32::try_from(info.size)
1079 .map_err(|_| Unsupported::Dynamic { inst, growing: Growing::Huge })?;
1080 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
1081
1082 // At least one, because the frame divides by the alignment and an object with no
1083 // alignment at all is one the front end had nothing to say about rather than one that may
1084 // go anywhere.
1085 let index = self.stack.locals.len();
1086 self.stack.locals.push(Local { size, align: info.align.max(1) });
1087
1088 let block = self.at.expect("a block is being filled");
1089 let reg = self.new_reg(result);
1090 let span = self.source.span(inst);
1091 let lea = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.lea)));
1092 let sp = mir::Operand::read(mir::Reg::physical(self.conv.stack_pointer), self.gpr);
1093 let made =
1094 self.out.build(block, lea).at(span).def(reg, self.gpr).mem(mir::Mem::at(sp)).finish();
1095 self.stack.addresses.push((made, index));
1096 Ok(())
1097 }
1098
1099 /// The other kind of `alloca`: one whose size the function does not know until it runs, which
1100 /// is what a variable length array is.
1101 ///
1102 /// Nothing about it is a slot the frame laid out, because the frame is laid out once and this
1103 /// happens as often as control reaches the declaration. The bytes come off the stack pointer
1104 /// where the declaration stands, which is two instructions:
1105 ///
1106 /// ```text
1107 /// sub sp, bytes the stack pointer moves down over the memory, which is what takes it
1108 /// lea reg, [sp+n] where the memory starts, which is above the outgoing argument area
1109 /// ```
1110 ///
1111 /// The displacement is left at nothing for the reason the constant kind leaves its own at
1112 /// nothing, and for a different number: that area belongs to the arguments of whatever this
1113 /// function calls, it stays at the bottom of the frame wherever the bottom has moved to, and
1114 /// how big it is is not known until every call in the function has been seen.
1115 ///
1116 /// The bytes are already a multiple of the stack pointer's alignment by the time they arrive,
1117 /// because [`crate::expand::rounds`] rounded them up in the IR, so nothing here has to mask the
1118 /// stack pointer afterwards and the stack pointer stays somewhere a call can be made from.
1119 ///
1120 /// Refused for an array wanting more alignment than the convention leaves the stack pointer
1121 /// with. Forcing that would be a second rounding of a register the frame already rounded, and
1122 /// after it no constant reaches the rest of the frame from anywhere. See `Growing` in
1123 /// [`crate::frame`].
1124 fn grow(&mut self, inst: Inst, size: Value) -> Result<(), Unsupported> {
1125 let data = &self.source[inst];
1126 let Extra::Mem(mem) = data.extra else { return Err(self.unsupported(inst)) };
1127 let info = self.source[mem];
1128 if info.align > self.conv.stack_align {
1129 return Err(Unsupported::Dynamic { inst, growing: Growing::Aligned });
1130 }
1131 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
1132 let bytes = self.reg_of(size)?;
1133
1134 let block = self.at.expect("a block is being filled");
1135 let span = self.source.span(inst);
1136 let stack = mir::Reg::physical(self.conv.stack_pointer);
1137 let grow = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.grow)));
1138 self.out
1139 .build(block, grow)
1140 .at(span)
1141 .operand(mir::Operand::write(stack, self.gpr))
1142 .operand(mir::Operand::read(stack, self.gpr))
1143 .operand(mir::Operand::read(bytes, self.gpr))
1144 .finish();
1145
1146 let reg = self.new_reg(result);
1147 let lea = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.lea)));
1148 let sp = mir::Operand::read(stack, self.gpr);
1149 let made =
1150 self.out.build(block, lea).at(span).def(reg, self.gpr).mem(mir::Mem::at(sp)).finish();
1151 self.stack.dynamic.push(made);
1152 self.stack.grown_at.get_or_insert(inst);
1153 Ok(())
1154 }
1155
1156 /// Where the stack pointer is, kept so that something later can put it back.
1157 ///
1158 /// One move out of the stack pointer and one move into it, which is the whole of what the two
1159 /// halves are. What makes them worth writing is where the front end puts them: a scope holding
1160 /// a variable length array saves the stack pointer as it opens and puts it back as it closes,
1161 /// so a loop declaring one takes its bytes once round rather than once per iteration, and a
1162 /// jump out of the scope gives the bytes back on the way out.
1163 ///
1164 /// The value travels in an ordinary register the allocator hands out, so it may be spilled like
1165 /// any other, and a spill slot in a frame that grows is reached through the frame pointer,
1166 /// which is exactly the register that still means something after the stack pointer has moved.
1167 fn stack_pointer(&mut self, inst: Inst, into: bool) -> Result<(), Unsupported> {
1168 let data = &self.source[inst];
1169 let block = self.at.expect("a block is being filled");
1170 let span = self.source.span(inst);
1171 let stack = mir::Reg::physical(self.conv.stack_pointer);
1172 let mov = x86_64::FRAME.moves(self.gpr).expect("a class the target says how to move").mov;
1173 let mov = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{mov}")));
1174 let (write, read) = if into {
1175 let &saved = self.source[data.args].first().ok_or_else(|| self.unsupported(inst))?;
1176 (stack, self.reg_of(saved)?)
1177 } else {
1178 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
1179 (self.new_reg(result), stack)
1180 };
1181 self.out
1182 .build(block, mov)
1183 .at(span)
1184 .operand(mir::Operand::write(write, self.gpr))
1185 .operand(mir::Operand::read(read, self.gpr))
1186 .finish();
1187 // Only the write is a move of the stack pointer, and it is the one that makes the frame a
1188 // growing one. A read of it in a function that never writes it back is a function that
1189 // asked where the stack was and did nothing with the answer.
1190 if into {
1191 self.stack.grown_at.get_or_insert(inst);
1192 }
1193 Ok(())
1194 }
1195
1196 /// Whether an instruction has an eighty bit float anywhere in it.
1197 ///
1198 /// Producing one and reading one are the same question here, because what makes one of these
1199 /// different from every other instruction is not the operation but where the value is. A
1200 /// `long double` is on the x87 stack while it is being worked on and in a frame slot the rest
1201 /// of the time, and neither of those is somewhere the operand of a rule could point.
1202 fn touches_x87(&self, inst: Inst) -> bool {
1203 let data = &self.source[inst];
1204 data.results().any(|value| on_x87(self.source[value].ty))
1205 || self.source[data.args].iter().any(|&arg| on_x87(self.source[arg].ty))
1206 }
1207
1208 /// Everything that happens to an eighty bit float, as the group of instructions it is.
1209 ///
1210 /// The first six move one, and every one of those is a load, a store, or a load and a store at
1211 /// two different formats, because that is the whole of what this machine converts with: the
1212 /// x87 has no instruction that turns one thing on its stack into another, so a widening is
1213 /// `fld` of the narrow format and a narrowing is `fstp` of it.
1214 ///
1215 /// The rest work on one, and they are here rather than in a rule for the same reason the six
1216 /// are. An add is a push, a push, the add and a pop, and what passes between those four is the
1217 /// top of a stack nothing allocates from, so there is no value in the middle of the group for
1218 /// a pattern to bind or a replacement to name. The comparison is the same shape with its last
1219 /// two instructions folded into one opcode, which is where the byte it produces comes from.
1220 ///
1221 /// Every group leaves the stack as empty as it found it, which is what `spec/10-backend.md`
1222 /// section 10.8 asks of one and is why nothing in this file has to track a depth: each push
1223 /// below is answered by a pop a line or two later, so no two groups can ever be looking at
1224 /// the same eight registers.
1225 fn x87(&mut self, inst: Inst) -> Result<(), Unsupported> {
1226 match self.source[inst].opcode {
1227 Opcode::Load => self.x87_load(inst),
1228 Opcode::Store => self.x87_store(inst),
1229 Opcode::FPExt => self.x87_widen(inst),
1230 Opcode::FPTrunc => self.x87_narrow(inst),
1231 Opcode::SIToFP => self.x87_from_signed(inst),
1232 Opcode::FPToSI => self.x87_to_signed(inst),
1233 Opcode::FAdd => self.x87_arith(inst, "fadd_p"),
1234 Opcode::FSub => self.x87_arith(inst, "fsubr_p"),
1235 Opcode::FMul => self.x87_arith(inst, "fmul_p"),
1236 Opcode::FDiv => self.x87_arith(inst, "fdivr_p"),
1237 Opcode::FNeg => self.x87_flip(inst),
1238 Opcode::FCmp => self.x87_compare(inst),
1239 Opcode::FConst => self.x87_const(inst),
1240 _ => Err(self.unsupported(inst)),
1241 }
1242 }
1243
1244 /// The eighty bit parameters of a block, copied out of the addresses an edge handed over and
1245 /// into slots of the block's own.
1246 ///
1247 /// What crosses an edge for a value of this type is an address, because the value is sixteen
1248 /// bytes of the frame and no register holds any of it. The block cannot keep that address: a
1249 /// second edge into the same block hands over a second one, and a read after the block would
1250 /// then be a read of whichever edge was taken rather than of one place. So the block has a
1251 /// slot per parameter and the bytes are copied into it here, which is the move on an edge that
1252 /// every other type gets from the allocator.
1253 ///
1254 /// Every load runs before every store and the stores run backwards, so all of the values are
1255 /// on the x87 stack at once and nothing reads a slot another one has already written. That
1256 /// costs nothing in the ordinary case of one parameter and is what makes the back edge of a
1257 /// loop that swaps two of these work. It is also the reason for the limit: the stack is eight
1258 /// deep, and a block with more of these than that is refused rather than copied in an order
1259 /// that could be wrong.
1260 fn settle(&mut self, block: Block, arriving: &[(Value, mir::Reg)]) -> Result<(), Unsupported> {
1261 let Some(&(first, _)) = arriving.first() else { return Ok(()) };
1262 if arriving.len() > X87_DEPTH {
1263 let ty = self.source[first].ty;
1264 return Err(Unsupported::Phi { block, count: arriving.len(), ty });
1265 }
1266 // A block parameter comes from no instruction, so what this points at is the first thing
1267 // in the block, which is where a reader looking for the copy would look.
1268 let first_inst = self.source.insts(block).next();
1269 let span = first_inst.map_or(Span::DUMMY, |it| self.source.span(it));
1270 for &(_, reg) in arriving {
1271 let from = self.through(reg);
1272 self.x87_at("fld_t", span, from);
1273 }
1274 for &(param, _) in arriving.iter().rev() {
1275 let into = self.x87_slot(param);
1276 let into = self.through(into);
1277 self.x87_at("fstp_t", span, into);
1278 }
1279 Ok(())
1280 }
1281
1282 /// The frame slot an eighty bit value lives in, as its address in a fresh register.
1283 ///
1284 /// The slot is the value's for the whole function and is taken the first time somebody asks.
1285 /// The address is worked out again every time, which is a `lea` per use and is deliberate: one
1286 /// address kept in a register from the definition to the last use would hold a general purpose
1287 /// register open across everything in between, and a function with a handful of these in it
1288 /// would spend its registers on addresses of things rather than on things.
1289 fn x87_slot(&mut self, value: Value) -> mir::Reg {
1290 // An argument of the function has a slot already and it is the caller's. The convention
1291 // puts the bytes in the argument area and hands over where they are, so the address that
1292 // arrived is the answer and no second copy of the value is made. Nothing ever writes to a
1293 // value of this type once it exists, so nothing writes to the caller's copy either. A
1294 // parameter of any other block is not this: what arrived there is an address a predecessor
1295 // chose, [`Lowering::settle`] has already copied the bytes out of it, and the slot those
1296 // bytes landed in is the one below.
1297 let entry = self.source.entry();
1298 if let (Def::Param { block, .. }, Some(reg)) =
1299 (self.source[value].def, self.regs[value.index()])
1300 {
1301 if entry == Some(block) {
1302 return reg;
1303 }
1304 }
1305 let index = match self.slots[value.index()] {
1306 Some(index) => index,
1307 None => {
1308 let index = self.stack.locals.len();
1309 self.stack.locals.push(Local { size: X87_BYTES, align: X87_BYTES });
1310 self.slots[value.index()] = Some(index);
1311 index
1312 }
1313 };
1314 let block = self.at.expect("a block is being filled");
1315 self.frame_address(block, index)
1316 }
1317
1318 /// The bytes a value crosses between a register and the x87 stack through, as their address
1319 /// in a fresh register.
1320 fn x87_crossing(&mut self) -> mir::Reg {
1321 let index = match self.crossing {
1322 Some(index) => index,
1323 None => {
1324 let index = self.stack.locals.len();
1325 self.stack.locals.push(Local { size: X87_CROSSING, align: X87_CROSSING });
1326 self.crossing = Some(index);
1327 index
1328 }
1329 };
1330 let block = self.at.expect("a block is being filled");
1331 self.frame_address(block, index)
1332 }
1333
1334 /// The two control words, as the address of the first of them in a fresh register.
1335 fn x87_control(&mut self) -> mir::Reg {
1336 let index = match self.control {
1337 Some(index) => index,
1338 None => {
1339 let index = self.stack.locals.len();
1340 self.stack.locals.push(Local { size: 4, align: 4 });
1341 self.control = Some(index);
1342 index
1343 }
1344 };
1345 let block = self.at.expect("a block is being filled");
1346 self.frame_address(block, index)
1347 }
1348
1349 /// An address held in a register, as the addressing mode that reaches it.
1350 fn through(&self, reg: mir::Reg) -> mir::Mem {
1351 mir::Mem::at(mir::Operand::read(reg, self.gpr))
1352 }
1353
1354 /// One instruction of a group, which names an address and nothing else.
1355 ///
1356 /// Every x87 instruction that moves a value is one of these. What it does to the stack is in
1357 /// the mnemonic rather than in an operand, so there is no register to write down and no
1358 /// register the allocator gets a say in.
1359 fn x87_at(&mut self, name: &str, span: Span, at: mir::Mem) {
1360 let block = self.at.expect("a block is being filled");
1361 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
1362 self.out.build(block, opcode).at(span).mem(at).finish();
1363 }
1364
1365 /// One instruction of a group that names nothing at all.
1366 ///
1367 /// The arithmetic is these. Both of an add's operands are already on the stack when it runs
1368 /// and so is where the answer goes, and the stack is not somewhere an instruction says, so
1369 /// `faddp` has an argument in the assembler's syntax and nothing here for the argument to come
1370 /// from. What it works on is which two pushes came before it, which is a fact about the order
1371 /// of the group and is why the group is written in one place.
1372 fn x87_only(&mut self, name: &str, span: Span) {
1373 let block = self.at.expect("a block is being filled");
1374 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
1375 self.out.build(block, opcode).at(span).finish();
1376 }
1377
1378 /// A `load` of a `long double`: onto the stack from where it was, and off it into the slot.
1379 ///
1380 /// Two instructions rather than the two general purpose moves the same sixteen bytes would
1381 /// take, because `fld` and `fstp` at this format neither convert nor look: the value goes on
1382 /// in the format it was already in and comes back off in it, so a signalling NaN stays one
1383 /// and nothing is raised. Which is what makes this a copy at all.
1384 fn x87_load(&mut self, inst: Inst) -> Result<(), Unsupported> {
1385 let (args, result) = self.ends(inst)?;
1386 let &address = args.first().ok_or_else(|| self.unsupported(inst))?;
1387 let span = self.source.span(inst);
1388 let from = self.reg_of(address)?;
1389 let from = self.through(from);
1390 let into = self.x87_slot(result);
1391 let into = self.through(into);
1392 self.x87_at("fld_t", span, from);
1393 self.x87_at("fstp_t", span, into);
1394 Ok(())
1395 }
1396
1397 /// A `store` of a `long double`: the same pair the other way round.
1398 fn x87_store(&mut self, inst: Inst) -> Result<(), Unsupported> {
1399 let args = self.source[self.source[inst].args].to_vec();
1400 let [value, address] = args[..] else { return Err(self.unsupported(inst)) };
1401 let span = self.source.span(inst);
1402 let from = self.x87_slot(value);
1403 let from = self.through(from);
1404 let into = self.reg_of(address)?;
1405 let into = self.through(into);
1406 self.x87_at("fld_t", span, from);
1407 self.x87_at("fstp_t", span, into);
1408 Ok(())
1409 }
1410
1411 /// A `float`, a `double` or an integer becoming a `long double`.
1412 ///
1413 /// Through memory, because the x87 reads memory and nothing else: the value is in a register
1414 /// the machine has and the unit has no way to be handed one, so it is written to the crossing
1415 /// bytes and loaded back at the format that widens it. Every one of these is exact. Sixty four
1416 /// bits of significand and fifteen of exponent hold every `float`, every `double` and every
1417 /// sixty four bit integer outright, so none of the four can round and none can raise.
1418 fn x87_across(
1419 &mut self,
1420 inst: Inst,
1421 put: &'static str,
1422 class: RegClass,
1423 get: &'static str,
1424 ) -> Result<(), Unsupported> {
1425 let (args, result) = self.ends(inst)?;
1426 let &source = args.first().ok_or_else(|| self.unsupported(inst))?;
1427 let span = self.source.span(inst);
1428 let value = self.reg_of(source)?;
1429 let across = self.x87_crossing();
1430 let across = self.through(across);
1431 let into = self.x87_slot(result);
1432 let into = self.through(into);
1433
1434 let block = self.at.expect("a block is being filled");
1435 let store = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{put}")));
1436 self.out.build(block, store).at(span).uses(value, class).mem(across).finish();
1437 self.x87_at(get, span, across);
1438 self.x87_at("fstp_t", span, into);
1439 Ok(())
1440 }
1441
1442 /// A `long double` becoming a `float`, a `double` or an integer.
1443 ///
1444 /// Through memory for the reason above and in the same three instructions backwards. The two
1445 /// that go to a float round to nearest, which is what the control word says unless somebody
1446 /// has changed it and is what C wants. The two that go to an integer do not, which is why they
1447 /// do not come here.
1448 fn x87_back(
1449 &mut self,
1450 inst: Inst,
1451 put: &'static str,
1452 get: &'static str,
1453 class: RegClass,
1454 ) -> Result<(), Unsupported> {
1455 let (args, result) = self.ends(inst)?;
1456 let &source = args.first().ok_or_else(|| self.unsupported(inst))?;
1457 let span = self.source.span(inst);
1458 let from = self.x87_slot(source);
1459 let from = self.through(from);
1460 let across = self.x87_crossing();
1461 let across = self.through(across);
1462
1463 self.x87_at("fld_t", span, from);
1464 self.x87_at(put, span, across);
1465 let block = self.at.expect("a block is being filled");
1466 let reg = self.new_reg(result);
1467 let load = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{get}")));
1468 self.out.build(block, load).at(span).def(reg, class).mem(across).finish();
1469 Ok(())
1470 }
1471
1472 /// An `fpext` up to a `long double`, which is the only direction this machine has one in.
1473 fn x87_widen(&mut self, inst: Inst) -> Result<(), Unsupported> {
1474 let sse = self.conv.sse_class;
1475 match self.source[self.narrow(inst)?].ty.bits() {
1476 32 => self.x87_across(inst, "movss_mr", sse, "fld_s"),
1477 64 => self.x87_across(inst, "movsd_mr", sse, "fld_l"),
1478 _ => Err(self.unsupported(inst)),
1479 }
1480 }
1481
1482 /// An `fptrunc` down from a `long double`, which is the other direction of the same.
1483 fn x87_narrow(&mut self, inst: Inst) -> Result<(), Unsupported> {
1484 let sse = self.conv.sse_class;
1485 let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
1486 match self.source[result].ty.bits() {
1487 32 => self.x87_back(inst, "fstp_s", "movss_rm", sse),
1488 64 => self.x87_back(inst, "fstp_l", "movsd_rm", sse),
1489 _ => Err(self.unsupported(inst)),
1490 }
1491 }
1492
1493 /// A `sitofp` up to a `long double`.
1494 ///
1495 /// Thirty two bits and sixty four, and nothing narrower, because C widens an integer to `int`
1496 /// before it converts one and the front end writes that widening down. An unsigned integer is
1497 /// not here at all: `fild` reads its operand as signed, so a value above the signed range
1498 /// comes back short by two to the sixty fourth and has to be added back, which is arithmetic
1499 /// rather than a move and waits with the rest of it.
1500 fn x87_from_signed(&mut self, inst: Inst) -> Result<(), Unsupported> {
1501 let gpr = self.gpr;
1502 match self.source[self.narrow(inst)?].ty.bits() {
1503 32 => self.x87_across(inst, "mov_mr_32", gpr, "fild_l"),
1504 64 => self.x87_across(inst, "mov_mr_64", gpr, "fild_ll"),
1505 _ => Err(self.unsupported(inst)),
1506 }
1507 }
1508
1509 /// An `fptosi` down from a `long double`, which is the one conversion here with no single
1510 /// instruction behind it.
1511 ///
1512 /// C cuts towards zero and the unit rounds the way its control word says, so the store that
1513 /// takes the value off the stack is wrapped in the control word being saved, changed and put
1514 /// back. Five instructions around the one that does the work, and three more moving the word
1515 /// through a register, because this machine has no way to OR a constant into memory at this
1516 /// width. The unit has a shorter answer in `fisttp`, and `spec/10-backend.md` section 10.8
1517 /// says why it is not used: it is SSE3, the x86-64 baseline is not, and there is nothing here
1518 /// that can gate an instruction on a feature yet.
1519 fn x87_to_signed(&mut self, inst: Inst) -> Result<(), Unsupported> {
1520 let (args, result) = self.ends(inst)?;
1521 let &source = args.first().ok_or_else(|| self.unsupported(inst))?;
1522 let (put, get) = match self.source[result].ty.bits() {
1523 32 => ("fistp_l", "mov_rm_32"),
1524 64 => ("fistp_ll", "mov_rm_64"),
1525 _ => return Err(self.unsupported(inst)),
1526 };
1527 let span = self.source.span(inst);
1528 let gpr = self.gpr;
1529 let from = self.x87_slot(source);
1530 let from = self.through(from);
1531 let across = self.x87_crossing();
1532 let across = self.through(across);
1533 let control = self.x87_control();
1534 let saved = self.through(control).plus(0);
1535 let cut = self.through(control).plus(2);
1536
1537 // The word the unit has now, into the first of the two slots and into a register, with the
1538 // rounding field turned to truncate on the way to the second.
1539 self.x87_at("fnstcw", span, saved);
1540 let block = self.at.expect("a block is being filled");
1541 let was = self.out.new_vreg(gpr);
1542 let read = mir::Opcode::new(self.names.intern("x64.mov_rm_16"));
1543 self.out.build(block, read).at(span).def(was, gpr).mem(saved).finish();
1544 let now = self.out.new_vreg(gpr);
1545 let set = mir::Opcode::new(self.names.intern("x64.or_ri_16"));
1546 // Two address, which is written out here rather than taken from the two shorthands
1547 // because the shorthands leave an operand unconstrained: this machine ORs into the
1548 // register it read, so the two have to be the same one and only the constraint says so.
1549 self.out
1550 .build(block, set)
1551 .at(span)
1552 .operand(mir::Operand::write(now, gpr).with(Constraint::Reuse(1)))
1553 .operand(mir::Operand::read(was, gpr))
1554 .imm(X87_TRUNCATE)
1555 .finish();
1556 let write = mir::Opcode::new(self.names.intern("x64.mov_mr_16"));
1557 self.out.build(block, write).at(span).uses(now, gpr).mem(cut).finish();
1558
1559 // The conversion itself, under the changed word, and then the word the unit had put back
1560 // before anything else runs.
1561 self.x87_at("fldcw", span, cut);
1562 self.x87_at("fld_t", span, from);
1563 self.x87_at(put, span, across);
1564 self.x87_at("fldcw", span, saved);
1565
1566 let block = self.at.expect("a block is being filled");
1567 let reg = self.new_reg(result);
1568 let load = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{get}")));
1569 self.out.build(block, load).at(span).def(reg, gpr).mem(across).finish();
1570 Ok(())
1571 }
1572
1573 /// A constant of this type, as the bits of it written into its slot.
1574 ///
1575 /// No x87 instruction at all, which is the surprise here. A slot holding an eighty bit value is
1576 /// the value, so a constant is ten bytes put where the value lives, and the unit never has to
1577 /// see it: whatever reads it will `fld` it out of the slot the way it reads any other one.
1578 ///
1579 /// Ten bytes in two goes, because the machine stores eight at a time and there is no store of
1580 /// an immediate to memory, so each half is put in a register first. The six bytes above the ten
1581 /// are left alone, since nothing reads them: they are the padding that makes the type sixteen
1582 /// wide and they are unspecified in the psABI rather than zero.
1583 ///
1584 /// The other way is a constant pool, an `fldt` of a symbol, and a relocation, which is what a
1585 /// compiler with somewhere to put a literal does. This back end has nowhere to put one yet, and
1586 /// four instructions in the frame is what that costs until it does.
1587 fn x87_const(&mut self, inst: Inst) -> Result<(), Unsupported> {
1588 let Extra::Imm(imm) = self.source[inst].extra else { return Err(self.unsupported(inst)) };
1589 let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
1590 let bits = self.source[imm].bits();
1591 let span = self.source.span(inst);
1592 let gpr = self.gpr;
1593 let slot = self.x87_slot(result);
1594 let low = self.through(slot).plus(0);
1595 let high = self.through(slot).plus(8);
1596
1597 let block = self.at.expect("a block is being filled");
1598 for (bytes, at, into) in
1599 [(bits as u64 as i64, low, "64"), (((bits >> 64) & 0xffff) as i64, high, "16")]
1600 {
1601 let held = self.out.new_vreg(gpr);
1602 let put = mir::Opcode::new(self.names.intern(&format!("{PREFIX}mov_ri_{into}")));
1603 self.out.build(block, put).at(span).def(held, gpr).imm(bytes).finish();
1604 let store = mir::Opcode::new(self.names.intern(&format!("{PREFIX}mov_mr_{into}")));
1605 self.out.build(block, store).at(span).uses(held, gpr).mem(at).finish();
1606 }
1607 Ok(())
1608 }
1609
1610 /// One arithmetic instruction on two eighty bit values, as the four it takes.
1611 ///
1612 /// The left operand is pushed first and the right one on top of it, so the left ends up
1613 /// underneath and the answer wanted is the one below against the top in that order. Which of
1614 /// the two mnemonics computes that is a question about the spelling rather than about the
1615 /// machine, and the two spellings disagree. Intel's `FSUBP ST(i), ST(0)` is `ST(i) - ST(0)`
1616 /// and is `DE E8+i`, and AT&T's `fsubp` is `DE E0+i`, which is the other subtraction. This
1617 /// compiler writes AT&T and encodes what gas encodes, so what it asks for here is `fsubr_p`
1618 /// and `fdivr_p`, and the `r` is not a reversal of anything the code generator decided.
1619 ///
1620 /// An addition and a multiplication have one form each and do not care, which is why a test
1621 /// that reads the mnemonic back would not have caught this and one that computes a subtraction
1622 /// and checks the answer does.
1623 ///
1624 /// The answer is left where the deeper of the two was and the shallower is gone, which is what
1625 /// the `p` on the mnemonic means, so one push has already been paid back by the time the
1626 /// `fstp` runs and the stack is level again after it.
1627 ///
1628 /// Nothing here is folded and nothing is reused. Two values that are the same value get two
1629 /// pushes of the same slot, and an operand that was just computed is read back out of the slot
1630 /// it was written to rather than left on the stack, which costs a store and a load per
1631 /// instruction in an expression. Keeping a partial result on the stack across the next
1632 /// instruction's operands means knowing how deep the stack is at every point in the block, and
1633 /// that is a different thing from writing a group.
1634 fn x87_arith(&mut self, inst: Inst, with: &'static str) -> Result<(), Unsupported> {
1635 let (args, result) = self.ends(inst)?;
1636 let [left, right] = args[..] else { return Err(self.unsupported(inst)) };
1637 let span = self.source.span(inst);
1638 let left = self.x87_slot(left);
1639 let left = self.through(left);
1640 let right = self.x87_slot(right);
1641 let right = self.through(right);
1642 let into = self.x87_slot(result);
1643 let into = self.through(into);
1644 self.x87_at("fld_t", span, left);
1645 self.x87_at("fld_t", span, right);
1646 self.x87_only(with, span);
1647 self.x87_at("fstp_t", span, into);
1648 Ok(())
1649 }
1650
1651 /// A negation, which is a push, the sign bit turned over and a pop.
1652 ///
1653 /// `fchs` does not read the value as a number, so this is right for a zero, for an infinity
1654 /// and for a NaN, and it raises nothing on any of them. Which is what C asks of a negation and
1655 /// is not what subtracting from zero would give: `0.0L - x` is a different answer at a
1656 /// negative zero and a signalling one at a NaN.
1657 fn x87_flip(&mut self, inst: Inst) -> Result<(), Unsupported> {
1658 let (args, result) = self.ends(inst)?;
1659 let &source = args.first().ok_or_else(|| self.unsupported(inst))?;
1660 let span = self.source.span(inst);
1661 let from = self.x87_slot(source);
1662 let from = self.through(from);
1663 let into = self.x87_slot(result);
1664 let into = self.through(into);
1665 self.x87_at("fld_t", span, from);
1666 self.x87_only("fchs", span);
1667 self.x87_at("fstp_t", span, into);
1668 Ok(())
1669 }
1670
1671 /// A comparison of two eighty bit values, as the two pushes and the one opcode that reads them.
1672 ///
1673 /// The right operand is pushed first and the left one on top of it, which is the other way
1674 /// round from the arithmetic and is because `fucomip` asks about the top against what is under
1675 /// it: the comparison this machine can do is the top's, so the value the predicate is about
1676 /// has to be the top. The pop that gets the loser off the stack and the byte that reads the
1677 /// flags are both inside the opcode, since what passes between those and the comparison is the
1678 /// flags and the flags are not something anything here can name.
1679 ///
1680 /// Which of the ten opcodes, and which way round, is the same table the vector comparisons
1681 /// match against in `rules/x86-64.rules`, and it has to stay the same table: a predicate that
1682 /// picked a different condition here than there would be a `long double` comparison that
1683 /// disagreed with the `double` comparison of the same two numbers, which is the one thing a
1684 /// wider format is not allowed to do.
1685 ///
1686 /// The always false and the always true are refused rather than folded into a constant,
1687 /// because a comparison this machine never has to do is one the optimizer should have removed
1688 /// and an instruction here that quietly agreed with it would hide that it did not.
1689 fn x87_compare(&mut self, inst: Inst) -> Result<(), Unsupported> {
1690 let Extra::FloatPred(pred) = self.source[inst].extra else {
1691 return Err(self.unsupported(inst));
1692 };
1693 let (args, result) = self.ends(inst)?;
1694 let [left, right] = args[..] else { return Err(self.unsupported(inst)) };
1695 // Two of the fourteen need a second byte and an instruction to put the two together,
1696 // because they are two conditions at once: an ordered equal is equal and not unordered,
1697 // and an unordered not equal is either. The opcode carries all of that and says here only
1698 // that it writes somewhere else as well.
1699 let (name, reversed, both) = match pred {
1700 FloatPred::Ogt => ("fucomip_set_a", false, false),
1701 FloatPred::Oge => ("fucomip_set_ae", false, false),
1702 FloatPred::Olt => ("fucomip_set_a", true, false),
1703 FloatPred::Ole => ("fucomip_set_ae", true, false),
1704 FloatPred::One => ("fucomip_set_ne", false, false),
1705 FloatPred::Ord => ("fucomip_set_np", false, false),
1706 FloatPred::Uno => ("fucomip_set_p", false, false),
1707 FloatPred::Ueq => ("fucomip_set_e", false, false),
1708 FloatPred::Ult => ("fucomip_set_b", false, false),
1709 FloatPred::Ule => ("fucomip_set_be", false, false),
1710 FloatPred::Ugt => ("fucomip_set_b", true, false),
1711 FloatPred::Uge => ("fucomip_set_be", true, false),
1712 FloatPred::Oeq => ("fucomip_set_e_and_np", false, true),
1713 FloatPred::Une => ("fucomip_set_ne_or_p", false, true),
1714 FloatPred::False | FloatPred::True => return Err(self.unsupported(inst)),
1715 };
1716 let (top, under) = if reversed { (right, left) } else { (left, right) };
1717
1718 let span = self.source.span(inst);
1719 let gpr = self.gpr;
1720 let under = self.x87_slot(under);
1721 let under = self.through(under);
1722 let top = self.x87_slot(top);
1723 let top = self.through(top);
1724 self.x87_at("fld_t", span, under);
1725 self.x87_at("fld_t", span, top);
1726
1727 let block = self.at.expect("a block is being filled");
1728 let reg = self.new_reg(result);
1729 // Taken before the instruction is started rather than inside it, since both come from the
1730 // same function being built and only one thing at a time may be adding to it.
1731 let spare = both.then(|| self.out.new_vreg(gpr));
1732 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
1733 let mut build = self.out.build(block, opcode).at(span).def(reg, gpr);
1734 if let Some(spare) = spare {
1735 build = build.def(spare, gpr);
1736 }
1737 build.finish();
1738 Ok(())
1739 }
1740
1741 /// The operands and the one result of an instruction that has exactly one.
1742 fn ends(&self, inst: Inst) -> Result<(&'a [Value], Value), Unsupported> {
1743 let data = &self.source[inst];
1744 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
1745 Ok((&self.source[data.args], result))
1746 }
1747
1748 /// The operand of a conversion, which is the end of it that is not the `long double`.
1749 fn narrow(&self, inst: Inst) -> Result<Value, Unsupported> {
1750 let args = &self.source[self.source[inst].args];
1751 args.first().copied().ok_or_else(|| self.unsupported(inst))
1752 }
1753
1754 /// One `va_start`, as the four fields of the list it was handed.
1755 ///
1756 /// Two of them are numbers this already knows, and each costs an instruction to put in a
1757 /// register before it can be stored, because the machine here has no store of an immediate to
1758 /// memory. The other two are addresses in the frame, and each is a `lea` [`crate::finish`]
1759 /// finishes: the save area is one of the function's own stack objects, and the caller's
1760 /// argument area is where the parameters that had no register came from, which is the same
1761 /// place and the same fixup a parameter past the sixth already uses.
1762 ///
1763 /// What is written is exactly the four fields [`crate::varargs`] describes, in the order they
1764 /// are laid out, so that reading this beside that table is the whole of the check.
1765 fn va_start(&mut self, inst: Inst) -> Result<(), Unsupported> {
1766 let Some(&list) = self.source[self.source[inst].args].first() else {
1767 return Err(self.unsupported(inst));
1768 };
1769 let started = self.varargs.ok_or_else(|| self.unsupported(inst))?;
1770 let list = self.reg_of(list)?;
1771 let block = self.at.expect("a block is being filled");
1772 let span = self.source.span(inst);
1773
1774 for (at, count) in
1775 [(varargs::GP_OFFSET, started.integers), (varargs::FP_OFFSET, started.floats)]
1776 {
1777 let held = self.out.new_vreg(self.gpr);
1778 let load = mir::Opcode::new(self.names.intern("x64.mov_ri_32"));
1779 self.out.build(block, load).at(span).def(held, self.gpr).imm(i64::from(count)).finish();
1780
1781 let store = mir::Opcode::new(self.names.intern("x64.mov_mr_32"));
1782 let mem = self.field(list, at);
1783 self.out.build(block, store).at(span).uses(held, self.gpr).mem(mem).finish();
1784 }
1785
1786 // The first argument the signature did not name, which is as far up the caller's argument
1787 // area as the ones it did name reached. Nothing here knows where that area is, so the
1788 // distance is recorded the way a parameter read out of it is and finished with it.
1789 let overflow = self.out.new_vreg(self.gpr);
1790 let lea = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.lea)));
1791 let sp = mir::Operand::read(mir::Reg::physical(self.conv.stack_pointer), self.gpr);
1792 let made = self
1793 .out
1794 .build(block, lea)
1795 .at(span)
1796 .def(overflow, self.gpr)
1797 .mem(mir::Mem::at(sp))
1798 .finish();
1799 self.stack.arguments.push((made, started.incoming));
1800
1801 let save = self.frame_address(block, started.save);
1802 for (at, held) in [(varargs::OVERFLOW, overflow), (varargs::SAVE_AREA, save)] {
1803 let store = mir::Opcode::new(self.names.intern("x64.mov_mr_64"));
1804 let mem = self.field(list, at);
1805 self.out.build(block, store).at(span).uses(held, self.gpr).mem(mem).finish();
1806 }
1807 Ok(())
1808 }
1809
1810 /// One field of a list, as the addressing mode that reaches it.
1811 fn field(&self, list: mir::Reg, at: i64) -> mir::Mem {
1812 let base = mir::Operand::read(list, self.gpr);
1813 mir::Mem::at(base).plus(i32::try_from(at).expect("a field of a list is a small offset"))
1814 }
1815
1816 /// The address of a name: one `lea` off the instruction pointer, with the name on it.
1817 ///
1818 /// The same instruction an `alloca` gets and for a related reason. An address that is not in
1819 /// the program is a `lea` of an addressing mode that names no register, and the mode carries
1820 /// the symbol so that [`rucc_asm`] can write it relative to `%rip` and leave the relocation
1821 /// for the assembler. Both halves of that already existed: the printer writes `sym(%rip)` and
1822 /// the encoder emits the relocation, because a call to a name the file does not define needed
1823 /// them first.
1824 ///
1825 /// One `mov` and not one `lea` when the name is one [`Elsewhere`] holds, because the distance
1826 /// the `lea` adds to the instruction pointer is a number only a link that puts the name in
1827 /// this program can work out, and the address of a function this file merely declares is not
1828 /// such a number. The load reads the address out of the slot the linker fills in instead. The
1829 /// linker turns it back into the `lea` when the name turns out to have been here all along,
1830 /// so this is not slower in the case that was already right.
1831 ///
1832 /// There is deliberately no name for this in [`crate::term`], which is what stops the address
1833 /// being folded into the instruction that reads it. Folding it is the right thing to do and
1834 /// is what turns a load of a global from two instructions into one, but it is a separate
1835 /// question about addressing modes and issue #282 is it. Until then the address is in a
1836 /// register before anything uses it, which is correct and one instruction longer.
1837 ///
1838 /// What this does not do is give the name anything to refer to. A module carries its globals
1839 /// and nothing writes them out, so a file that defines the variable it reads compiles to a
1840 /// reference the linker cannot resolve. Issue #293 is the other half.
1841 ///
1842 /// A thread-local variable is neither of the two above and is [`Self::thread_address`].
1843 fn address_of(&mut self, inst: Inst) -> Result<(), Unsupported> {
1844 let data = &self.source[inst];
1845 let Extra::Symbol(symbol) = data.extra else { return Err(self.unsupported(inst)) };
1846 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
1847 if self.elsewhere.thread(symbol) {
1848 return self.thread_address(inst, symbol, result);
1849 }
1850
1851 let block = self.at.expect("a block is being filled");
1852 let reg = self.new_reg(result);
1853 let span = self.source.span(inst);
1854 let (mnemonic, mem) = if self.elsewhere.holds(symbol) {
1855 (GOT_LOAD, mir::Mem::got(symbol))
1856 } else {
1857 (x86_64::FRAME.lea, mir::Mem::of(symbol))
1858 };
1859 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{mnemonic}")));
1860 self.out.build(block, opcode).at(span).def(reg, self.gpr).mem(mem).finish();
1861 Ok(())
1862 }
1863
1864 /// The address of a thread-local variable, which is this thread's copy of it.
1865 ///
1866 /// Neither instruction the ordinary case writes would mean anything here. There is no distance
1867 /// to the variable for a `lea` to add, because there is no variable: there is one copy of it per
1868 /// thread and they are at different addresses, so a link asked for the distance to the name
1869 /// refuses rather than picking one. And there is no address for a table slot to hold either, for
1870 /// the same reason.
1871 ///
1872 /// What is the same in every thread is where the variable sits inside the block of storage a
1873 /// thread gets, so that offset is what the link writes down, and the address of the running
1874 /// thread's block is what turns it into an address. x86-64 keeps that address in `%fs`, at the
1875 /// front of the block, so the whole of this is three instructions:
1876 ///
1877 /// ```text
1878 /// movq x@gottpoff(%rip), %off # how far into the block x sits, which the link fills in
1879 /// movq %fs:0, %tp # where this thread's block is, which only the machine knows
1880 /// addq %tp, %off # this thread's copy of x
1881 /// ```
1882 ///
1883 /// That is the initial exec model. It is one instruction longer than what gcc writes at `-O2`
1884 /// in an executable, which folds the addition into the instruction that uses the address, and
1885 /// the difference is issue #282 rather than anything about threads: nothing here folds an
1886 /// address into its reader yet. The link relaxes the first instruction into an immediate when it
1887 /// is making an executable, since it lays the blocks out and therefore knows the number, so the
1888 /// table slot costs nothing in the case that is common.
1889 ///
1890 /// It is not the most general model. A library loaded by `dlopen` gets its storage after the
1891 /// program is already running, and the block this reaches was laid out before it started, so
1892 /// the loader has to find room in that block for the library's variables. glibc keeps a little
1893 /// spare room for exactly this and a library that fits in it loads and runs; one that does not
1894 /// fails to load, with a message saying so. The model with no such limit calls `__tls_get_addr`
1895 /// and is what gcc writes under `-fPIC` by default, and it is issue #1104.
1896 ///
1897 /// So this is the model gcc writes under `-ftls-model=initial-exec`: right for an executable,
1898 /// right for a library the program is linked against, and a load that either works or is
1899 /// refused out loud for a library something opens later. What it is never is quietly wrong.
1900 fn thread_address(
1901 &mut self,
1902 inst: Inst,
1903 symbol: Symbol,
1904 result: Value,
1905 ) -> Result<(), Unsupported> {
1906 let block = self.at.expect("a block is being filled");
1907 let span = self.source.span(inst);
1908 let gpr = self.gpr;
1909 let load = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{GOT_LOAD}")));
1910
1911 let offset = self.out.new_vreg(gpr);
1912 self.out
1913 .build(block, load)
1914 .at(span)
1915 .def(offset, gpr)
1916 .mem(mir::Mem::thread(symbol))
1917 .finish();
1918 // The front of the block, which is the one thing on this machine that no instruction can
1919 // work out: `%fs` is not a register a program can read, and what it points at is a word
1920 // holding its own address, so reading through it at zero is how the address is come by.
1921 let pointer = self.out.new_vreg(gpr);
1922 let at = mir::Mem::in_segment(Segment::Fs, 0);
1923 self.out.build(block, load).at(span).def(pointer, gpr).mem(at).finish();
1924
1925 // Two address, spelled out for the reason `x87_to_int` gives: this machine adds into the
1926 // register it read, and only the constraint says the two are the same one.
1927 let reg = self.new_reg(result);
1928 let add = mir::Opcode::new(self.names.intern(&format!("{PREFIX}add_rr_64")));
1929 self.out
1930 .build(block, add)
1931 .at(span)
1932 .operand(mir::Operand::write(reg, gpr).with(Constraint::Reuse(1)))
1933 .operand(mir::Operand::read(offset, gpr))
1934 .operand(mir::Operand::read(pointer, gpr))
1935 .finish();
1936 Ok(())
1937 }
1938
1939 /// `__builtin_thread_pointer`, which is the front of the block [`Self::thread_address`] adds
1940 /// an offset to.
1941 ///
1942 /// The same one instruction, on its own this time and with nothing to add to it. A program
1943 /// writes this when what it wants is a number that is different in every thread and cheap to
1944 /// come by, rather than a variable of its own in the block, so there is no relocation here and
1945 /// no name for the link to resolve.
1946 fn thread_pointer(&mut self, inst: Inst) -> Result<(), Unsupported> {
1947 let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
1948 let block = self.at.expect("a block is being filled");
1949 let span = self.source.span(inst);
1950 let reg = self.new_reg(result);
1951 let load = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{GOT_LOAD}")));
1952 let at = mir::Mem::in_segment(Segment::Fs, 0);
1953 self.out.build(block, load).at(span).def(reg, self.gpr).mem(at).finish();
1954 Ok(())
1955 }
1956
1957 /// A conversion that converts nothing: the result is the operand under another type.
1958 ///
1959 /// `ptrtoint` and `inttoptr` at one width are the whole of this. An address on this machine is
1960 /// an integer as wide as the machine addresses, so a cast between the two changes what the
1961 /// type system calls the value and changes nothing about the value, and the register holding
1962 /// it is the register that already held it. The front end never writes either of them at any
1963 /// other width, because it widens or narrows around the cast rather than through it, so the
1964 /// two widths disagreeing here means the IR came from somewhere else and is refused rather
1965 /// than guessed at.
1966 ///
1967 /// Reading the operand first is what materializes it when it is a constant, which is the case
1968 /// that matters: a null pointer is an `inttoptr` of zero, and that zero has to reach a
1969 /// register before anything can call it an address.
1970 fn rename(&mut self, inst: Inst) -> Result<(), Unsupported> {
1971 let data = &self.source[inst];
1972 let [arg] = self.source[data.args] else { return Err(self.unsupported(inst)) };
1973 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
1974 if !self.is_address_width(self.source[arg].ty)
1975 || !self.is_address_width(self.source[result].ty)
1976 {
1977 return Err(self.unsupported(inst));
1978 }
1979 let reg = self.reg_of(arg)?;
1980 self.regs[result.index()] = Some(reg);
1981 Ok(())
1982 }
1983
1984 /// One barrier, which on this machine is one instruction at the strongest ordering and no
1985 /// instruction at all at every other one.
1986 ///
1987 /// x86-64 is total store order, so the only reordering the machine does is a store followed by
1988 /// a load of a different address, and the only ordering that forbids that is sequential
1989 /// consistency. An acquire, a release and an acquire release fence are therefore already true
1990 /// of every program running here, and what a program wanted from writing one is that the
1991 /// compiler not move memory accesses across it. The optimizer has finished by the time this
1992 /// runs and nothing below reorders one access past another, so the constraint is already
1993 /// discharged and there is nothing to write.
1994 ///
1995 /// The strongest one is `mfence`, which is what gcc 16.2.0 writes for
1996 /// `__atomic_thread_fence(__ATOMIC_SEQ_CST)` and for `__sync_synchronize`. A locked instruction
1997 /// on the stack is faster on most parts and is what some compilers write instead; it is also a
1998 /// write to memory the program did not ask for, and the plain barrier is the one that says what
1999 /// it means.
2000 ///
2001 /// Written here by name rather than by a rule, for the same reason a `lea` of a symbol is:
2002 /// there is nothing in a barrier that a proof over bitvectors could discharge. It computes
2003 /// nothing, so there is no equality to state, and what makes it the right answer is the memory
2004 /// model, which the rule language cannot talk about.
2005 fn barrier(&mut self, inst: Inst) -> Result<(), Unsupported> {
2006 let Extra::Order(order) = self.source[inst].extra else {
2007 return Err(self.unsupported(inst));
2008 };
2009 if order != MemOrder::SeqCst {
2010 return Ok(());
2011 }
2012 let block = self.at.expect("a block is being filled");
2013 let span = self.source.span(inst);
2014 let fence = mir::Opcode::new(self.names.intern("x64.mfence"));
2015 self.out.build(block, fence).at(span).finish();
2016 Ok(())
2017 }
2018
2019 /// One hint that an address is about to be used, which is one instruction and no promise.
2020 ///
2021 /// Four instructions on this machine and the locality picks between them, which is what the
2022 /// number means: how much of the data will still be wanted after the access. None of it wanted
2023 /// is `prefetchnta`, which brings the line in without keeping it, and all of it wanted is
2024 /// `prefetcht0`, which brings it as close as the machine can. The two in between are the levels
2025 /// between those. Measured against gcc 16.2.0 on x86-64 rather than read off the manual: zero
2026 /// gives `prefetchnta`, one `prefetcht2`, two `prefetcht1` and three `prefetcht0`.
2027 ///
2028 /// Whether the access will write is not read here, and that is this machine rather than an
2029 /// omission. The write hint is `prefetchw`, which is not in the base instruction set, and gcc
2030 /// writes it only when the command line said the part has it. So a prefetch for a write is the
2031 /// same instruction as a prefetch for a read, which is what gcc 16.2.0 writes without
2032 /// `-mprfchw`, and the difference is carried in the IR for a target that can use it.
2033 ///
2034 /// The address goes in the addressing mode rather than in an operand, the way a store's does.
2035 /// It is built here as the plainest one there is, a register and nothing else, because what
2036 /// arrives is a value and folding an addition into the mode is a rule's job and no rule reaches
2037 /// this instruction. An address the program computed is therefore one `lea` or one add in front
2038 /// of this, which is what it would have been for the load the hint is about anyway.
2039 fn hint(&mut self, inst: Inst) -> Result<(), Unsupported> {
2040 let Extra::Prefetch(hint) = self.source[inst].extra else {
2041 return Err(self.unsupported(inst));
2042 };
2043 let args: Vec<Value> = self.source[self.source[inst].args].to_vec();
2044 let [address] = args[..] else { return Err(self.unsupported(inst)) };
2045 let name = match hint.locality {
2046 0 => "prefetch_nta",
2047 1 => "prefetch_t2",
2048 2 => "prefetch_t1",
2049 PrefetchHint::MOST => "prefetch_t0",
2050 // Nothing else exists. The checker reads a locality outside the range as zero and the
2051 // verifier refuses one that got here another way, so this is a hint that was built
2052 // rather than checked, and the safe answer for a hint is to write no instruction.
2053 _ => return Err(self.unsupported(inst)),
2054 };
2055 let base = self.reg_of(address)?;
2056 let block = self.at.expect("a block is being filled");
2057 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
2058 self.out
2059 .build(block, opcode)
2060 .at(self.source.span(inst))
2061 .mem(mir::Mem::at(mir::Operand::read(base, self.gpr)))
2062 .finish();
2063 Ok(())
2064 }
2065
2066 /// One compare and exchange, which is the instruction every other atomic on this machine is
2067 /// built out of.
2068 ///
2069 /// What the IR asks for is: read what is at an address, compare it against a value the program
2070 /// expected, put a second value there if the two were equal, and say both what was read and
2071 /// whether the exchange happened. The machine has exactly that instruction, and the `lock` in
2072 /// front of it is what makes the whole of it one step as far as every other processor is
2073 /// concerned.
2074 ///
2075 /// The ordering is not read here, and that is the memory model rather than an omission. A
2076 /// locked instruction on x86-64 is a full barrier whatever the program asked for, so a relaxed
2077 /// compare and exchange and a sequentially consistent one are the same instruction, and there
2078 /// is nothing weaker to emit for the weaker orderings. The failure ordering is not read for the
2079 /// same reason.
2080 ///
2081 /// The two values it produces are why this is written by name. The one the program compares
2082 /// against and the one it gets back are both `rax`, which the instruction reads and writes
2083 /// without being told, and the table says so with a fixed constraint at each end rather than
2084 /// leaving the allocator to find out. The second value is the byte behind it, which is the zero
2085 /// flag read out by a `sete`, and it is a definition of the same instruction so that the
2086 /// allocator knows the two are live together and never gives the byte the register the answer
2087 /// is in.
2088 fn exchange(&mut self, inst: Inst) -> Result<(), Unsupported> {
2089 let args: Vec<Value> = self.source[self.source[inst].args].to_vec();
2090 let results: Vec<Value> = self.source[inst].results().collect();
2091 let [addr, expected, desired] = args[..] else { return Err(self.unsupported(inst)) };
2092 let [old, exchanged] = results[..] else { return Err(self.unsupported(inst)) };
2093
2094 // A value the machine can compare in one instruction, which is an integer or an address at
2095 // one of the four widths it has a compare and exchange for. Anything else is a type this
2096 // has no instruction for rather than a program that is wrong, and the front end refuses it
2097 // before ever getting here.
2098 let ty = self.source[old].ty;
2099 let bits = if ty.is_ptr() { ADDRESS_BITS } else { ty.bits() };
2100 if (!ty.is_int() && !ty.is_ptr()) || !matches!(bits, 8 | 16 | 32 | 64) {
2101 return Err(self.unsupported(inst));
2102 }
2103
2104 let base = self.reg_of(addr)?;
2105 let want = self.reg_of(expected)?;
2106 let put = self.reg_of(desired)?;
2107 let got = self.new_reg(old);
2108 let flag = self.new_reg(exchanged);
2109
2110 let name = format!("cmpxchg_{bits}");
2111 let form = x86_64::form(&name).ok_or_else(|| self.unsupported(inst))?;
2112 let block = self.at.expect("a block is being filled");
2113 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
2114 let mut build = self.out.build(block, opcode).at(self.source.span(inst));
2115 for (desc, reg) in form.operands().iter().zip([got, flag, want, put]) {
2116 let operand = mir::Operand {
2117 reg,
2118 class: desc.class,
2119 role: desc.role,
2120 constraint: desc.constraint,
2121 };
2122 build = build.operand(operand);
2123 }
2124 build.mem(mir::Mem::at(mir::Operand::read(base, self.gpr))).finish();
2125 Ok(())
2126 }
2127
2128 /// One read modify write, for the three operations this machine does in a single instruction.
2129 ///
2130 /// What the IR asks for is: read what is at an address, do something to it, put the answer back,
2131 /// say what was there before, and let nothing get between the three steps. The machine has
2132 /// `xchg` for putting a value there and `lock xadd` for adding one, and both leave what they
2133 /// found in the register the operand arrived in, which is why the value that comes back and the
2134 /// value that went in are one register here.
2135 ///
2136 /// A subtraction is the add over the negated operand, which is right at every width because the
2137 /// machine's arithmetic wraps and negating then adding is subtracting in two's complement
2138 /// whatever the operands were. The negate is a separate instruction in front, over a register of
2139 /// its own, so that the value the program handed over is not the one written on: an operand may
2140 /// be live after this and a program that read it again would read the negation.
2141 ///
2142 /// The ordering is not read, for the reason the compare and exchange beside this does not read
2143 /// it. `xchg` with memory locks the bus whether it is asked to or not and `lock xadd` is asked
2144 /// to, so both are full barriers on this machine and there is nothing weaker to fall to.
2145 ///
2146 /// Eight of the other ten never arrive, because `crate::retry` turned each of them into a loop
2147 /// around a compare and exchange before anything here saw it. The two that do arrive are the
2148 /// ones on floating values, and they are refused: a compare and exchange of a float wants the
2149 /// value carried through an integer of the same width, and an eighty bit float has no such
2150 /// width. Neither family of builtins can write one yet either, so a program that reaches this
2151 /// refusal is a program that reached an unimplemented builtin first.
2152 fn modify(&mut self, inst: Inst) -> Result<(), Unsupported> {
2153 let Extra::Rmw(op, _) = self.source[inst].extra else {
2154 return Err(self.unsupported(inst));
2155 };
2156 let args: Vec<Value> = self.source[self.source[inst].args].to_vec();
2157 let [addr, operand] = args[..] else { return Err(self.unsupported(inst)) };
2158 let old = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
2159
2160 // A value the machine can exchange in one instruction, which is an integer at one of the
2161 // four widths it has these for. A pointer arrives as an address, so it is an integer by the
2162 // time it is here, and anything else is a type this has no instruction for.
2163 let ty = self.source[old].ty;
2164 if !ty.is_int() || !matches!(ty.bits(), 8 | 16 | 32 | 64) {
2165 return Err(self.unsupported(inst));
2166 }
2167 let name = match op {
2168 RmwOp::Xchg => format!("xchg_{}", ty.bits()),
2169 RmwOp::Add | RmwOp::Sub => format!("xadd_{}", ty.bits()),
2170 _ => return Err(self.unsupported(inst)),
2171 };
2172
2173 let base = self.reg_of(addr)?;
2174 let mut put = self.reg_of(operand)?;
2175 let block = self.at.expect("a block is being filled");
2176 let span = self.source.span(inst);
2177 if op == RmwOp::Sub {
2178 let negated = self.out.new_vreg(self.gpr);
2179 let negate =
2180 mir::Opcode::new(self.names.intern(&format!("{PREFIX}neg_r_{}", ty.bits())));
2181 let form = x86_64::form(&format!("neg_r_{}", ty.bits()))
2182 .ok_or_else(|| self.unsupported(inst))?;
2183 let mut build = self.out.build(block, negate).at(span);
2184 for (desc, reg) in form.operands().iter().zip([negated, put]) {
2185 build = build.operand(mir::Operand {
2186 reg,
2187 class: desc.class,
2188 role: desc.role,
2189 constraint: desc.constraint,
2190 });
2191 }
2192 build.finish();
2193 put = negated;
2194 }
2195
2196 let got = self.new_reg(old);
2197 let form = x86_64::form(&name).ok_or_else(|| self.unsupported(inst))?;
2198 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
2199 let mut build = self.out.build(block, opcode).at(span);
2200 for (desc, reg) in form.operands().iter().zip([got, put]) {
2201 build = build.operand(mir::Operand {
2202 reg,
2203 class: desc.class,
2204 role: desc.role,
2205 constraint: desc.constraint,
2206 });
2207 }
2208 build.mem(mir::Mem::at(mir::Operand::read(base, self.gpr))).finish();
2209 Ok(())
2210 }
2211
2212 /// One `asm` statement, for as long as its template has no instructions in it.
2213 ///
2214 /// An empty template is most of the inline assembly in a test suite, and it is not a corner
2215 /// case somebody wrote by accident. A program that wants a value computed where it stands, or a
2216 /// loop the optimizer must not touch, writes `asm volatile ("" : : : "memory")`, and forty
2217 /// years of bug reports about optimizers are full of them. What such a statement asks for is
2218 /// the barrier and the operand places, and no instructions at all.
2219 ///
2220 /// So the instructions are the easy half here and there are none of them. The half that is
2221 /// real is the operands: a constraint says where a value has to be, and where it has to be is
2222 /// still true when the template between them is empty.
2223 ///
2224 /// What the constraints ask for, on an empty template, is only ever that two operands share a
2225 /// place. Nothing reads a register no text names, so `"r"` on its own asks for a register and
2226 /// no particular one, and any register at all answers it. A matching constraint is different,
2227 /// because it says the output the assembly leaves is the place the input arrived in, and with
2228 /// no instructions between them that is the input unchanged. So it is a rename and not a move:
2229 /// the value is already in a register and the result is that register.
2230 ///
2231 /// An output nothing is tied to is whatever the assembly left there, which for a template that
2232 /// writes nothing is whatever was in the register. That is a value the program is not entitled
2233 /// to, and this writes a zero rather than reading one, because the allocator has to be given a
2234 /// definition before a use whatever the program is entitled to.
2235 ///
2236 /// The clobber list is not read, and on an empty template that is right rather than an
2237 /// omission. A clobber says the assembly ruins a register, and a template with no instructions
2238 /// in it ruins nothing.
2239 fn assembly(&mut self, inst: Inst) -> Result<(), Unsupported> {
2240 let data = &self.source[inst];
2241 let Extra::Asm(asm) = data.extra else { return Err(self.unsupported(inst)) };
2242 let info = self.source[asm];
2243 if !self.source[info.targets].is_empty() {
2244 return Err(Unsupported::Assembly { inst, refused: Written::Goto });
2245 }
2246 if !self.names.resolve(info.template).trim().is_empty() {
2247 return Err(Unsupported::Assembly { inst, refused: Written::Template });
2248 }
2249
2250 let constraints = self.names.resolve(info.constraints).to_string();
2251 let results: Vec<Value> = data.results().collect();
2252 let operands = AsmOperands::read(&constraints, &results, &self.source[data.args])
2253 .ok_or(Unsupported::Assembly { inst, refused: Written::Operand })?;
2254
2255 for (index, operand) in operands.iter().copied().enumerate().collect::<Vec<_>>() {
2256 let Some(result) = operand.result else { continue };
2257 let ty = self.source[result].ty;
2258 if on_x87(ty) {
2259 return Err(Unsupported::Assembly { inst, refused: Written::Operand });
2260 }
2261 match operands.tied_to(index) {
2262 // The place the input arrived in, which the assembly wrote nothing over.
2263 Some(from) => {
2264 if self.class_of(self.source[from].ty) != self.class_of(ty) {
2265 return Err(Unsupported::Assembly { inst, refused: Written::Operand });
2266 }
2267 let reg = self.reg_of(from)?;
2268 self.regs[result.index()] = Some(reg);
2269 }
2270 None => self.undefined(inst, result)?,
2271 }
2272 }
2273 Ok(())
2274 }
2275
2276 /// A register holding a value the program has no claim on, written as a zero.
2277 ///
2278 /// Every other way of saying it costs the same instruction or needs a word the machine IR does
2279 /// not have, and a zero is the one that reads the same on every run.
2280 fn undefined(&mut self, inst: Inst, result: Value) -> Result<(), Unsupported> {
2281 let ty = self.source[result].ty;
2282 let refused = Unsupported::Assembly { inst, refused: Written::Operand };
2283 if self.class_of(ty) != self.gpr || !matches!(ty.bits(), 8 | 16 | 32 | 64) {
2284 return Err(refused);
2285 }
2286 let block = self.at.expect("a block is being filled");
2287 let span = self.source.span(inst);
2288 let reg = self.new_reg(result);
2289 let put = mir::Opcode::new(self.names.intern(&format!("{PREFIX}mov_ri_{}", ty.bits())));
2290 self.out.build(block, put).at(span).def(reg, self.gpr).imm(0).finish();
2291 Ok(())
2292 }
2293
2294 /// Whether a type is the width an address is, which is what makes a cast to or from one free.
2295 fn is_address_width(&self, ty: Type) -> bool {
2296 ty.is_ptr() || (ty.is_int() && ty.bits() == ADDRESS_BITS)
2297 }
2298
2299 /// Where a block goes, which in machine IR is on the block rather than on its terminator.
2300 ///
2301 /// That is why no rule ever names a block: a branch is selected for what it reads and the
2302 /// edges are copied across here, arguments and all. The arguments are read last, after every
2303 /// instruction of the block is written, because an argument that is a constant is
2304 /// materialized where it is first wanted and the end of the block is where an edge wants it.
2305 ///
2306 /// Which is not quite the end. A block that leaves two ways has the branch as its last
2307 /// instruction, and anything appended after a branch is something the branch has already
2308 /// jumped past, so a constant materialized here would be a register the block below reads and
2309 /// nothing ever writes. The branch is put back on the end when that happened, which is the
2310 /// only reordering anything in this crate does and is why the branch is remembered before a
2311 /// single argument is read.
2312 fn edges(&mut self, block: Block, out: mir::Block) -> Result<(), Unsupported> {
2313 let Some(term) = self.source.terminator(block) else { return Ok(()) };
2314 let branch =
2315 if self.source[term].opcode == Opcode::BrIf { self.out.terminator(out) } else { None };
2316
2317 let calls: Vec<rucc_ir::BlockCall> = self.source.successors(term).collect();
2318 let mut succs = Vec::with_capacity(calls.len());
2319 for call in calls {
2320 let args: Vec<Value> = self.source[call.args].to_vec();
2321 let mut regs = Vec::with_capacity(args.len());
2322 for value in args {
2323 // The address of where the value is rather than the value, for the one type a
2324 // register holds none of. The block on the other side copies the bytes out of it
2325 // into a slot of its own, which is what makes a second edge into the same block
2326 // safe.
2327 let reg = if on_x87(self.source[value].ty) {
2328 self.x87_slot(value)
2329 } else {
2330 self.reg_of(value)?
2331 };
2332 regs.push(reg);
2333 }
2334 succs.push(mir::BlockCall::with(self.out_block(call.block), regs));
2335 }
2336 if let Some(branch) = branch {
2337 if self.out.terminator(out) != Some(branch) {
2338 self.out.remove_inst(branch);
2339 self.out.append_inst(out, branch);
2340 }
2341 }
2342 *self.out.succs_mut(out) = succs;
2343 Ok(())
2344 }
2345
2346 /// The machine IR block an IR block became.
2347 fn out_block(&self, block: Block) -> mir::Block {
2348 self.blocks[block.index()].expect("every block was created before any was filled")
2349 }
2350
2351 /// The parameters of the entry block, which are the function's arguments.
2352 ///
2353 /// They are not block parameters in the machine IR and they cannot be. A block parameter is
2354 /// given its value by a move on the edge into the block, and there is no edge into an entry
2355 /// block, so what arrives in a function is the convention's to say. [`crate::abi`] is what
2356 /// says it.
2357 ///
2358 /// The ones past the last register arrived in the caller's memory and are read out of it, and
2359 /// the loads that read them come back here so that the frame can finish them the way it
2360 /// finishes an `alloca`.
2361 fn arrive(&mut self, block: Block, out: mir::Block) -> Result<(), Unsupported> {
2362 let params = self.source[block].params.clone();
2363 // The type of each is the block's answer and what the ABI asks of it is the signature's,
2364 // and the two lists are the same list: a parameter the classification turned into a
2365 // pointer is a pointer in the block too. A block with more parameters than the signature
2366 // names is not one the front end writes, and each of those is taken as a plain value.
2367 let asked: Vec<Abi> = self.source.signature().params.iter().map(|it| it.abi).collect();
2368 let types: Vec<Param> = params
2369 .iter()
2370 .enumerate()
2371 .map(|(index, &value)| {
2372 let abi = asked.get(index).copied().unwrap_or_default();
2373 Param { ty: self.source[value].ty, abi }
2374 })
2375 .collect();
2376 // A save area for a function that takes arguments its signature does not name, on a
2377 // convention whose list is the four field one. Windows is the other kind and has no area at
2378 // all, so a `va_start` in one is refused rather than built wrong.
2379 let variadic = self.source.signature().variadic && !self.conv.shared_positions;
2380 let area = variadic.then(|| varargs::Area::of(self.conv));
2381 let arrived = abi::entry(&mut self.out, out, &types, self.conv, self.names, area)
2382 .map_err(|(index, missing)| Unsupported::Argument { index, missing })?;
2383 for (¶m, reg) in params.iter().zip(&arrived.regs) {
2384 self.regs[param.index()] = Some(*reg);
2385 }
2386 if let Some(area) = area {
2387 self.save_area(out, &arrived, area);
2388 }
2389 self.stack.arguments.extend(arrived.stack);
2390 Ok(())
2391 }
2392
2393 /// The prologue of a variadic function, which is every argument register it was handed written
2394 /// into the frame.
2395 ///
2396 /// Every one the signature did not name, that is. Which of those hold anything is a thing only
2397 /// the caller knew and there is nothing here to ask, so all of them are written, and the ones a
2398 /// named parameter took are not, because `va_start` sets the two offsets past them and nothing
2399 /// ever reads their slots.
2400 ///
2401 /// What that costs is up to fourteen stores in the prologue of a function that may read none of
2402 /// them, and the convention's answer to that is the count of vector registers in `%al`, which
2403 /// lets a callee skip the eight vector stores when the call passed no floats. Skipping them is a
2404 /// branch in a prologue, and a prologue is written long after this by [`crate::finish`], which
2405 /// has no blocks to branch between. So they are all written every time, which is correct and is
2406 /// what `-O0` costs. Issue #323 is the branch.
2407 ///
2408 /// A vector register is written eight bytes at a time and not sixteen, for the reason
2409 /// [`crate::varargs`] gives: the upper half of a slot is not something any reader of a list
2410 /// looks at.
2411 ///
2412 /// The address is computed once into a register rather than written as a displacement off the
2413 /// stack pointer, because a displacement into a frame is not known until after allocation and
2414 /// one `lea` costs less than a fixup list for a dozen stores. It is the same `lea` an `alloca`
2415 /// gets and [`crate::finish`] fills it in the same way.
2416 fn save_area(&mut self, out: mir::Block, arrived: &abi::Arrived, area: varargs::Area) {
2417 let save = self.stack.locals.len();
2418 self.stack.locals.push(Local { size: area.size, align: varargs::VECTOR_SLOT });
2419 self.varargs = Some(Varargs {
2420 save,
2421 incoming: arrived.used,
2422 integers: u32::try_from(arrived.took.0).unwrap_or(0) * area.stride(false),
2423 floats: area.starts_at(true)
2424 + u32::try_from(arrived.took.1).unwrap_or(0) * area.stride(true),
2425 });
2426
2427 let base = self.frame_address(out, save);
2428 for &(reg, class, at) in &arrived.spare {
2429 let name = if class == self.gpr { "x64.mov_mr_64" } else { "x64.movsd_mr" };
2430 let store = mir::Opcode::new(self.names.intern(name));
2431 let up = i32::try_from(at).expect("a register save area under two gigabytes");
2432 let mem = mir::Mem::at(mir::Operand::read(base, self.gpr)).plus(up);
2433 self.out.build(out, store).uses(reg, class).mem(mem).finish();
2434 }
2435 }
2436
2437 /// The address of one of the function's stack objects, in a fresh register.
2438 ///
2439 /// Written with nothing in its displacement, because where an object is in a frame is not known
2440 /// until after allocation, and given to [`crate::finish`] to fill in the way an `alloca` is.
2441 fn frame_address(&mut self, out: mir::Block, local: usize) -> mir::Reg {
2442 let reg = self.out.new_vreg(self.gpr);
2443 let lea = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.lea)));
2444 let sp = mir::Operand::read(mir::Reg::physical(self.conv.stack_pointer), self.gpr);
2445 let made = self.out.build(out, lea).def(reg, self.gpr).mem(mir::Mem::at(sp)).finish();
2446 self.stack.addresses.push((made, local));
2447 reg
2448 }
2449
2450 /// Whether an instruction is one no machine instruction is written for where it stands.
2451 ///
2452 /// Four of them, and none is a lowering decision, which is why none is a rule. A constant is
2453 /// written where a register for it is first wanted rather than where the IR put it, and every
2454 /// reader of one may have folded it into an immediate, in which case nowhere is the right
2455 /// place. A return of nothing has nothing to put anywhere: the epilogue gives the frame back
2456 /// and leaves, and it is appended to every block with no successors long after this has
2457 /// finished, so a return with a value is one instruction here and a return without one is
2458 /// none. Unless the value went back through memory, in which case there is something to put
2459 /// somewhere after all and the IR does not carry it: the address the caller handed over has
2460 /// to be in `rax` on the way out, and [`Lowering::returned`] is what writes that.
2461 ///
2462 /// An unconditional jump is the third, and there is even less of it: the edge is on the
2463 /// block, and whether the block it goes to is the next one and needs no jump at all is the
2464 /// block layout's answer rather than this one's.
2465 ///
2466 /// The fourth is a point control does not arrive at, in both of the forms the IR has for it:
2467 /// the `unreachable` terminator the front end puts at the end of a function whose body can run
2468 /// off the bottom, and the `unreachable_hint` a call to `__builtin_unreachable` becomes. What
2469 /// to write for a place nothing reaches is a question with no wrong answer, and nothing is the
2470 /// smallest one and the one gcc 16.2.0 gives at `-O0`. The terminator leaves the block with no
2471 /// successors, so the epilogue lands at the end of it the way it does on any other block that
2472 /// goes nowhere, and the function cannot fall out of its own last instruction into whatever
2473 /// the assembler puts next.
2474 fn writes_nothing(&self, inst: Inst) -> bool {
2475 let data = &self.source[inst];
2476 match data.opcode {
2477 Opcode::IConst | Opcode::Jump | Opcode::Unreachable | Opcode::UnreachableHint => true,
2478 Opcode::Return => self.source[data.args].is_empty() && self.sret().is_none(),
2479 _ => false,
2480 }
2481 }
2482
2483 /// What every instruction in one block matched, with a set of values nobody may take.
2484 ///
2485 /// Backwards, because an instruction that has been folded into a later one does not get to
2486 /// fold anything into itself: the rule that took it only reached one level down, so what is
2487 /// under it is not in the term the matcher saw and cannot be replaced.
2488 fn decide(&self, insts: &[Inst], refused: &HashSet<Value>) -> Decided {
2489 let mut found: Vec<Option<Match<Term>>> = (0..insts.len()).map(|_| None).collect();
2490 let mut plans: Vec<Option<Plan>> = vec![None; insts.len()];
2491 let mut folded: Vec<Inst> = Vec::new();
2492 for (index, &inst) in insts.iter().enumerate().rev() {
2493 if folded.contains(&inst) {
2494 continue;
2495 }
2496 if let Some((plan, matched)) = self.select(inst, refused) {
2497 folded.extend(self.folds(inst, plan));
2498 found[index] = Some(matched);
2499 plans[index] = Some(plan);
2500 }
2501 }
2502 Decided { found, plans, folded }
2503 }
2504
2505 /// A value some of its readers took and some of them did not, which is the one case folding
2506 /// buys nothing.
2507 ///
2508 /// Folding does not delete the instruction that computed a value for anybody else, so a
2509 /// reader that did not take it still needs it in a register and the instruction stays. The
2510 /// reader that did take it now does that work again. Either all of them take it, in which
2511 /// case nothing is left to read it and the instruction goes, or none of them do.
2512 ///
2513 /// The count is over the whole function rather than over the block, since a value read from
2514 /// another block is read from a register there whatever this block decides. An instruction
2515 /// built by name rather than matched, a call being the one that matters, has no plan and so
2516 /// takes nothing, which is the right answer for it as well.
2517 fn left_alive(&self, insts: &[Inst], plans: &[Option<Plan>]) -> Option<Value> {
2518 let mut taken = vec![0u32; self.uses.len()];
2519 for (&inst, plan) in insts.iter().zip(plans) {
2520 let Some(plan) = plan else { continue };
2521 let args = &self.source[self.source[inst].args];
2522 for (index, &arg) in args.iter().take(MAX_ARGS).enumerate() {
2523 if plan[index] == Shown::Expand {
2524 taken[arg.index()] += 1;
2525 }
2526 }
2527 }
2528 for (&inst, plan) in insts.iter().zip(plans) {
2529 let Some(plan) = plan else { continue };
2530 let args = &self.source[self.source[inst].args];
2531 for (index, &arg) in args.iter().take(MAX_ARGS).enumerate() {
2532 if plan[index] == Shown::Expand && taken[arg.index()] < self.uses[arg.index()] {
2533 return Some(arg);
2534 }
2535 }
2536 }
2537 None
2538 }
2539
2540 /// The rule that fires on an instruction, and what it bound.
2541 ///
2542 /// The plans are tried in order and the first that matches wins, which is the maximal munch
2543 /// `spec/10-backend.md` asks for: a plan that offers more to the matcher is tried before one
2544 /// that offers less.
2545 fn select(&self, inst: Inst, refused: &HashSet<Value>) -> Option<(Plan, Match<Term>)> {
2546 for plan in self.plans(inst, refused) {
2547 let terms = Terms::new(self.source, inst, plan);
2548 if let Some(matched) = TABLE.find(&terms, Term::Root) {
2549 return Some((plan, matched));
2550 }
2551 }
2552 None
2553 }
2554
2555 /// Every way this instruction can be shown to the matcher, most offered first.
2556 fn plans(&self, inst: Inst, refused: &HashSet<Value>) -> Vec<Plan> {
2557 let args = &self.source[self.source[inst].args];
2558 let mut plans = vec![PLAIN];
2559 for (index, &arg) in args.iter().enumerate().take(MAX_ARGS) {
2560 let mut ways = Vec::new();
2561 if self.foldable(inst, arg, refused) {
2562 ways.push(Shown::Expand);
2563 }
2564 if Terms::new(self.source, inst, PLAIN).constant(arg).is_some() {
2565 ways.push(Shown::Const);
2566 }
2567 ways.push(Shown::Reg);
2568 plans = plans
2569 .into_iter()
2570 .flat_map(|plan| {
2571 ways.iter().map(move |&way| {
2572 let mut next = plan;
2573 next[index] = way;
2574 next
2575 })
2576 })
2577 .collect();
2578 }
2579 plans
2580 }
2581
2582 /// Whether an operand may be shown as the instruction that computed it.
2583 ///
2584 /// It has to be in the same block, because a rule that folds one instruction into another
2585 /// moves the work to where the second one is. It has to be something rather than a block
2586 /// parameter, and not a constant, which is shown as a constant instead. And it has to be a
2587 /// value [`Lowering::left_alive`] has not put back, which is how the one reader at a time
2588 /// question is asked here: this says yes to a value with any number of readers, and a value
2589 /// only some of them could take is refused after the fact and asked again.
2590 ///
2591 /// A value with several readers used to be refused outright, on the reasoning that folding
2592 /// does not delete the instruction for anybody else. That reasoning is about the set of
2593 /// readers and was being applied to one reader at a time, which is stricter than it needs to
2594 /// be: when every reader takes it there is nobody left to read it and the instruction goes.
2595 /// An address a store and a load share is the shape that matters, since a memory operand has
2596 /// room for the whole of it and both readers have a memory operand.
2597 fn foldable(&self, into: Inst, value: Value, refused: &HashSet<Value>) -> bool {
2598 let Def::Result { inst, .. } = self.source[value].def else { return false };
2599 if self.source[inst].opcode == Opcode::IConst || refused.contains(&value) {
2600 return false;
2601 }
2602 self.source.block_of(inst).is_some()
2603 && self.source.block_of(inst) == self.source.block_of(into)
2604 }
2605
2606 /// The instructions a match folded into the one it matched.
2607 ///
2608 /// The plan is what says this, not the bindings: a binding is a register or a number either
2609 /// way, and an operand shown as the instruction that computed it is one no rule could have
2610 /// matched without taking that instruction, because the plan offered the matcher nothing
2611 /// else to call it.
2612 fn folds(&self, inst: Inst, plan: Plan) -> Vec<Inst> {
2613 let args = &self.source[self.source[inst].args];
2614 args.iter()
2615 .take(MAX_ARGS)
2616 .enumerate()
2617 .filter(|&(index, _)| plan[index] == Shown::Expand)
2618 .filter_map(|(_, &arg)| match self.source[arg].def {
2619 Def::Result { inst, .. } => Some(inst),
2620 Def::Param { .. } => None,
2621 })
2622 .collect()
2623 }
2624
2625 /// Build the machine instruction a match calls for.
2626 fn emit(&mut self, inst: Inst, matched: &Match<Term>) -> Result<(), Unsupported> {
2627 let rule: &Rule = TABLE.rule(matched);
2628 let pieces = rule.replacement;
2629 let Some(Piece::App { head, arity }) = pieces.first() else {
2630 return Err(self.unsupported(inst));
2631 };
2632 let opcode = head.strip_prefix(PREFIX).ok_or_else(|| self.unsupported(inst))?;
2633 let form = x86_64::form(opcode).ok_or_else(|| self.unsupported(inst))?;
2634
2635 let mut read = Read::default();
2636 let mut at = 1;
2637 for _ in 0..*arity {
2638 at = self.read(inst, pieces, at, &matched.bindings, &mut read)?;
2639 }
2640
2641 let descs = form.operands();
2642 let writes = descs.iter().take_while(|desc| desc.role.is_def()).count();
2643 if descs.len() - writes != read.regs.len() {
2644 return Err(self.unsupported(inst));
2645 }
2646
2647 // The first thing the instruction writes is what it computes, and any others are
2648 // registers the machine destroys on the way, which are fresh because nothing else is in
2649 // them and nothing reads them. An instruction that writes nothing at all is one whose
2650 // whole purpose is its effect, which is what a store is, and there is no result to put
2651 // anywhere.
2652 let mut regs = Vec::new();
2653 if writes > 0 {
2654 let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
2655 regs.push(self.new_reg(result));
2656 // The rest are the registers the machine destroys on the way, and the class each is in
2657 // is the one the instruction's description gives it rather than a guess, so that an
2658 // instruction that wrecks a register in the other file says so.
2659 regs.extend(descs[1..writes].iter().map(|desc| self.out.new_vreg(desc.class)));
2660 } else if self.source[inst].first_result.is_some() {
2661 // A rule that throws away a value the IR gave a name to would leave every reader of
2662 // that name with nothing to read, so it is a rule this and the target disagree about.
2663 return Err(self.unsupported(inst));
2664 }
2665 regs.extend(read.regs.iter().copied());
2666
2667 let block = self.at.expect("a block is being filled");
2668 let opcode = mir::Opcode::new(self.names.intern(head));
2669 let mut build = self.out.build(block, opcode).at(self.source.span(inst));
2670 for (desc, reg) in descs.iter().zip(regs) {
2671 let operand = mir::Operand {
2672 reg,
2673 class: desc.class,
2674 role: desc.role,
2675 constraint: desc.constraint,
2676 };
2677 build = build.operand(operand);
2678 }
2679 if let Some(mem) = read.mem {
2680 build = build.mem(mem);
2681 }
2682 if let Some(imm) = read.imm {
2683 build = build.imm(imm);
2684 }
2685 build.finish();
2686 Ok(())
2687 }
2688
2689 /// Read one argument of a replacement, which is a register, a number or an address.
2690 ///
2691 /// Gives back the position after it, because a replacement is flat and an address takes
2692 /// arguments of its own.
2693 fn read(
2694 &mut self,
2695 inst: Inst,
2696 pieces: &'static [Piece],
2697 at: usize,
2698 bindings: &[Term],
2699 out: &mut Read,
2700 ) -> Result<usize, Unsupported> {
2701 match pieces.get(at) {
2702 Some(Piece::Int(value)) => {
2703 out.imm = i64::try_from(*value).ok();
2704 Ok(at + 1)
2705 }
2706 Some(Piece::Var { index, .. }) => {
2707 match bindings.get(*index) {
2708 Some(&Term::Reg(value)) => {
2709 let reg = self.reg_of(value)?;
2710 out.regs.push(reg);
2711 }
2712 Some(&Term::Num(value)) => out.imm = i64::try_from(value).ok(),
2713 // A pattern binds a register or a number and nothing else, so this is a
2714 // rule the matcher and this file disagree about.
2715 _ => return Err(self.unsupported(inst)),
2716 }
2717 Ok(at + 1)
2718 }
2719 Some(Piece::App { head, arity }) => {
2720 let kind = x86_64::address(head).ok_or_else(|| self.unsupported(inst))?;
2721 let mut inner = Read::default();
2722 let mut next = at + 1;
2723 for _ in 0..*arity {
2724 next = self.read(inst, pieces, next, bindings, &mut inner)?;
2725 }
2726 let mem = address(kind, &inner, self.gpr).ok_or_else(|| self.unsupported(inst))?;
2727 out.mem = Some(mem);
2728 Ok(next)
2729 }
2730 None => Err(self.unsupported(inst)),
2731 }
2732 }
2733
2734 /// The register a value is in, materializing it if it is a constant that has not been put in
2735 /// one yet.
2736 ///
2737 /// A constant is written where it is wanted rather than where the IR defined it, and where it
2738 /// is wanted is a block that need not be the one the IR defined it in. So the register holding
2739 /// one is only good inside the block it was written into, and a second block that wants the
2740 /// same constant gets its own. Anything else is a register read where nothing wrote it: the
2741 /// IR guarantees a definition dominates its uses, and this moved the definition.
2742 ///
2743 /// Writing the number again is also the right answer and not merely the safe one. It is one
2744 /// instruction that reads nothing, which is cheaper than holding a register live across a
2745 /// branch for it, and it is what a rematerializing allocator would do with the value anyway.
2746 fn reg_of(&mut self, value: Value) -> Result<mir::Reg, Unsupported> {
2747 let constant = match self.source[value].def {
2748 Def::Result { inst, .. } => {
2749 (self.source[inst].opcode == Opcode::IConst).then_some(inst)
2750 }
2751 Def::Param { .. } => None,
2752 };
2753 let here = self.at.expect("a block is being filled");
2754 if let Some(reg) = self.regs[value.index()] {
2755 if constant.is_none() || self.written[value.index()] == Some(here) {
2756 return Ok(reg);
2757 }
2758 }
2759 if let Some(inst) = constant {
2760 // Cleared so that the register the constant is written into is a new one rather than
2761 // the one the block above wrote, which is still being read up there.
2762 self.regs[value.index()] = None;
2763 // Nothing is refused here. A constant is written on its own, out of the loop over the
2764 // block, and the operands of the rule that writes one are the number and nothing else.
2765 let matched = self
2766 .select(inst, &HashSet::new())
2767 .map(|(_, matched)| matched)
2768 .ok_or_else(|| self.unsupported(inst))?;
2769 self.emit(inst, &matched)?;
2770 // The same mark the loop over the instructions makes, and it has to be made here as
2771 // well because this is the only place a constant is ever selected: the loop skips one
2772 // where the IR wrote it, so a rule that lowers a constant fires from nowhere else and
2773 // would be reported as a rule nothing reaches.
2774 self.fired.mark(matched.rule);
2775 self.written[value.index()] = Some(here);
2776 return Ok(self.regs[value.index()].expect("a constant is written into a register"));
2777 }
2778 Ok(self.new_reg(value))
2779 }
2780
2781 /// Which register file a value of that type lives in.
2782 ///
2783 /// The vector one for the two float widths the machine has scalar instructions for, and the
2784 /// general purpose one for everything else. A `long double` is in neither, and it is here
2785 /// rather than in the vector class on purpose: it would be put in a register that cannot hold
2786 /// it, and there is no rule that names one, so the instruction computing it is reported. The
2787 /// wrong class would make that a wrong program instead of a refused one.
2788 fn class_of(&self, ty: Type) -> RegClass {
2789 match crate::term::float_slot(ty) {
2790 Some(_) => self.conv.sse_class,
2791 None => self.gpr,
2792 }
2793 }
2794
2795 /// A fresh register for a value, which is what the instruction computing it writes.
2796 fn new_reg(&mut self, value: Value) -> mir::Reg {
2797 if let Some(reg) = self.regs[value.index()] {
2798 return reg;
2799 }
2800 let reg = self.out.new_vreg(self.class_of(self.source[value].ty));
2801 self.regs[value.index()] = Some(reg);
2802 reg
2803 }
2804
2805 fn unsupported(&self, inst: Inst) -> Unsupported {
2806 let data = &self.source[inst];
2807 Unsupported::Inst {
2808 inst,
2809 term: Terms::new(self.source, inst, PLAIN).name(inst),
2810 opcode: data.opcode,
2811 ty: data.first_result.map(|result| self.source[result].ty),
2812 }
2813 }
2814}
2815
2816/// What the arguments of one replacement came to.
2817#[derive(Debug, Default)]
2818struct Read {
2819 regs: Vec<mir::Reg>,
2820 imm: Option<i64>,
2821 mem: Option<mir::Mem>,
2822}
2823
2824/// The addressing mode an address constructor's arguments make.
2825///
2826/// One arm per constructor rather than a question asked of the kind, because what the arguments
2827/// mean is the whole of what tells the four apart: the same register is a base in one and an
2828/// index in another, and the same constant is a scale in one and a displacement in another.
2829fn address(kind: x86_64::Address, read: &Read, gpr: RegClass) -> Option<mir::Mem> {
2830 let mut regs = read.regs.iter().copied().map(|reg| mir::Operand::read(reg, gpr));
2831 match kind {
2832 x86_64::Address::BaseIndexScale => {
2833 let base = regs.next()?;
2834 let index = regs.next()?;
2835 Some(mir::Mem::at(base).indexed(index, u8::try_from(read.imm?).ok()?))
2836 }
2837 x86_64::Address::IndexScale => Some(mir::Mem {
2838 base: None,
2839 index: Some(regs.next()?),
2840 scale: u8::try_from(read.imm?).ok()?,
2841 disp: 0,
2842 symbol: None,
2843 reach: mir::Reach::Itself,
2844 segment: None,
2845 }),
2846 x86_64::Address::Base => Some(mir::Mem::at(regs.next()?)),
2847 // The rule that writes this has a guard saying the constant fits, so a displacement that
2848 // does not is a rule and a target that disagree rather than a program this cannot compile.
2849 x86_64::Address::BaseOffset => {
2850 Some(mir::Mem { disp: i32::try_from(read.imm?).ok()?, ..mir::Mem::at(regs.next()?) })
2851 }
2852 }
2853}
2854
2855/// The table this selector matches with.
2856///
2857/// One target for now, because one target has a rule file. Which table to use becomes a question
2858/// the moment a second one does, and the answer will be the target the session was given rather
2859/// than a constant here.
2860static TABLE: &Table = &crate::select::x86_64::TABLE;
2861
2862#[cfg(test)]
2863mod tests {
2864 use rucc_ir::{
2865 AsmInfo, Builder, CallInfo, Flags, InstData, MemInfo, MemOrder, Restrict, Signature, Type,
2866 };
2867 use rucc_regalloc::assign::Env;
2868 use rucc_target::x86_64::{FRAME, REGS, SYSV};
2869
2870 use super::*;
2871 use crate::finish::{Convention, finish};
2872 use crate::frame::{Frame, Incoming, Layout};
2873
2874 /// A function of as many 64 bit parameters as the test wants, and the block they are in.
2875 fn blank(params: &[Type]) -> (Interner, Func, Block, Vec<Value>) {
2876 let mut names = Interner::new();
2877 let mut func = Func::new(names.intern("f"), Signature::new());
2878 let block = func.create_block();
2879 let values = params.iter().map(|&ty| func.append_param(block, ty)).collect();
2880 (names, func, block, values)
2881 }
2882
2883 /// An ordinary access: not atomic, and aligned enough that nothing here has an opinion.
2884 /// Neither field reaches selection, which is the point of saying it once here.
2885 fn plain() -> MemInfo {
2886 MemInfo {
2887 size: 0,
2888 align: 1,
2889 order: MemOrder::NotAtomic,
2890 tbaa: None,
2891 owns: 0,
2892 restrict: Restrict::NONE,
2893 }
2894 }
2895
2896 /// What the allocator is given: every integer register the convention offers except two, held
2897 /// back so that a move on an edge has somewhere to break a cycle and a spilled value has
2898 /// somewhere to be read into. Which two does not matter, and holding back the last two the
2899 /// convention would reach for leaves every expectation below unchanged.
2900 fn env() -> Env {
2901 const SCRATCH: [rucc_target::PhysReg; 2] = [x86_64::R10, x86_64::R11];
2902 let order: Vec<rucc_target::PhysReg> =
2903 SYSV.int_order.iter().copied().filter(|reg| !SCRATCH.contains(reg)).collect();
2904 Env::new().with(x86_64::GPR, &order, &SCRATCH)
2905 }
2906
2907 /// The machine IR text a function lowers to.
2908 fn lower(names: &mut Interner, source: &Func) -> String {
2909 let out = func(source, names, &SYSV, &Elsewhere::default())
2910 .expect("every instruction has a rule");
2911 mir::print_func(&out.func, names, ®S)
2912 }
2913
2914 #[test]
2915 fn an_addition_of_two_registers_is_one_instruction() {
2916 let i32 = Type::int(32);
2917 let (mut names, mut func, block, args) = blank(&[i32, i32]);
2918 let mut build = Builder::new(&mut func, block);
2919 build.binary(Opcode::Add, args[0], args[1], Flags::default());
2920
2921 assert_eq!(
2922 lower(&mut names, &func),
2923 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
2924 %1:gpr($rsi) = x64.arg_val_32\n %2:gpr(reuse 1) = x64.add_rr_32 %0, %1\n}\n"
2925 );
2926 }
2927
2928 #[test]
2929 fn a_constant_operand_becomes_an_immediate() {
2930 let i32 = Type::int(32);
2931 let (mut names, mut func, block, args) = blank(&[i32]);
2932 let mut build = Builder::new(&mut func, block);
2933 let seven = build.iconst(i32, 7);
2934 build.binary(Opcode::Add, args[0], seven, Flags::default());
2935
2936 // The constant is in the instruction and nothing was written to hold it, which is what
2937 // materializing one where a register for it is wanted buys.
2938 assert_eq!(
2939 lower(&mut names, &func),
2940 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
2941 %1:gpr(reuse 1) = x64.add_ri_32 %0, 7\n}\n"
2942 );
2943 }
2944
2945 #[test]
2946 fn a_constant_too_wide_for_an_immediate_goes_into_a_register() {
2947 let i64 = Type::int(64);
2948 let (mut names, mut func, block, args) = blank(&[i64]);
2949 let mut build = Builder::new(&mut func, block);
2950 let big = build.iconst(i64, i128::from(i32::MAX) + 1);
2951 build.binary(Opcode::Add, args[0], big, Flags::default());
2952
2953 // Nobody wrote this fallback down. The rule that takes an immediate has a guard that
2954 // turns a number this wide down, so it does not fire, and the next way of showing the
2955 // operand puts it in a register.
2956 assert_eq!(
2957 lower(&mut names, &func),
2958 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
2959 %1:gpr = x64.mov_ri_64 2147483648\n %2:gpr(reuse 1) = x64.add_rr_64 %0, %1\n}\n"
2960 );
2961 }
2962
2963 #[test]
2964 fn an_index_calculation_folds_into_an_address() {
2965 let i64 = Type::int(64);
2966 let (mut names, mut func, block, args) = blank(&[i64, i64]);
2967 let mut build = Builder::new(&mut func, block);
2968 let four = build.iconst(i64, 4);
2969 let scaled = build.binary(Opcode::Mul, args[1], four, Flags::default());
2970 build.binary(Opcode::Add, args[0], scaled, Flags::default());
2971
2972 // Three IR instructions and one machine instruction. The multiply is gone because the
2973 // rule that matched reached down and took it.
2974 assert_eq!(
2975 lower(&mut names, &func),
2976 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
2977 %1:gpr($rsi) = x64.arg_val_64\n %2:gpr = x64.lea_64 [%0 + %1*4]\n}\n"
2978 );
2979 }
2980
2981 #[test]
2982 fn an_instruction_every_reader_can_take_is_folded_into_all_of_them() {
2983 let i64 = Type::int(64);
2984 let (mut names, mut func, block, args) = blank(&[i64, i64]);
2985 let mut build = Builder::new(&mut func, block);
2986 let four = build.iconst(i64, 4);
2987 let scaled = build.binary(Opcode::Mul, args[1], four, Flags::default());
2988 let first = build.binary(Opcode::Add, args[0], scaled, Flags::default());
2989 build.binary(Opcode::Add, first, scaled, Flags::default());
2990
2991 // Both readers have room for a scaled index, so both of them take it and nothing is left
2992 // to read the multiply. Three IR instructions become two machine ones, where refusing to
2993 // fold into either reader would have left three.
2994 assert_eq!(
2995 lower(&mut names, &func),
2996 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
2997 %1:gpr($rsi) = x64.arg_val_64\n %2:gpr = x64.lea_64 [%0 + %1*4]\n \
2998 %3:gpr = x64.lea_64 [%2 + %1*4]\n}\n"
2999 );
3000 }
3001
3002 #[test]
3003 fn an_instruction_one_of_its_readers_cannot_take_is_folded_into_none_of_them() {
3004 let i64 = Type::int(64);
3005 let (mut names, mut func, block, args) = blank(&[i64, i64]);
3006 let mut build = Builder::new(&mut func, block);
3007 let four = build.iconst(i64, 4);
3008 let scaled = build.binary(Opcode::Mul, args[1], four, Flags::default());
3009 build.binary(Opcode::Add, args[0], scaled, Flags::default());
3010 build.store(scaled, args[0], plain(), Flags::default());
3011
3012 // The addition has room for the multiply and the store does not: what a store writes is
3013 // a register, and no rule reaches through it. Folding into the addition alone would
3014 // leave the multiply where it is for the store to read and do the work twice, so the
3015 // multiply is put back and both readers read the register it wrote.
3016 let text = lower(&mut names, &func);
3017 assert!(text.contains("x64.lea_64 [%1*4]"), "{text}");
3018 assert!(text.contains("x64.add_rr_64"), "{text}");
3019 }
3020
3021 #[test]
3022 fn a_shift_by_a_register_asks_for_it_in_cl() {
3023 let i32 = Type::int(32);
3024 let (mut names, mut func, block, args) = blank(&[i32, i32]);
3025 let mut build = Builder::new(&mut func, block);
3026 build.binary(Opcode::Shl, args[0], args[1], Flags::default());
3027
3028 // The fixed register is not in the rule. It is what the target says the instruction does
3029 // with its operands, and the allocator is what will act on it.
3030 let text = lower(&mut names, &func);
3031 assert!(text.contains("x64.shl_rcl_32 %0, %1($rcx)"), "{text}");
3032 }
3033
3034 #[test]
3035 fn a_division_names_the_registers_and_the_register_it_destroys() {
3036 let i32 = Type::int(32);
3037 let (mut names, mut func, block, args) = blank(&[i32, i32]);
3038 let mut build = Builder::new(&mut func, block);
3039 build.binary(Opcode::SDiv, args[0], args[1], Flags::default());
3040
3041 // Two definitions, because a division writes the remainder whether anybody wanted it or
3042 // not, and the second one is early because it is destroyed before the operands are read.
3043 let text = lower(&mut names, &func);
3044 assert!(
3045 text.contains("%2:gpr($rax), early %3:gpr($rdx) = x64.idiv_quo_32 %0($rax), %1"),
3046 "{text}"
3047 );
3048 }
3049
3050 #[test]
3051 fn a_load_reads_through_the_register_the_address_is_in() {
3052 let i64 = Type::int(64);
3053 let (mut names, mut func, block, args) = blank(&[i64]);
3054 let mut build = Builder::new(&mut func, block);
3055 build.load(Type::int(32), args[0], plain(), Flags::default());
3056
3057 assert_eq!(
3058 lower(&mut names, &func),
3059 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
3060 %1:gpr = x64.mov_rm_32 [%0]\n}\n"
3061 );
3062 }
3063
3064 #[test]
3065 fn a_store_writes_no_register_and_the_value_it_writes_is_the_one_the_ir_gave_it() {
3066 let (mut names, mut func, block, args) = blank(&[Type::int(32), Type::int(64)]);
3067 let mut build = Builder::new(&mut func, block);
3068 build.store(args[0], args[1], plain(), Flags::default());
3069
3070 // The value is the first parameter and the address is the second, and the instruction
3071 // takes them the other way round. Getting that backwards would compile to a store of the
3072 // address into the value, which is a program that runs and does the wrong thing.
3073 assert_eq!(
3074 lower(&mut names, &func),
3075 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
3076 %1:gpr($rsi) = x64.arg_val_64\n x64.mov_mr_32 %0, [%1]\n}\n"
3077 );
3078 }
3079
3080 #[test]
3081 fn an_address_with_a_constant_added_folds_into_the_access() {
3082 let i64 = Type::int(64);
3083 let (mut names, mut func, block, args) = blank(&[i64]);
3084 let mut build = Builder::new(&mut func, block);
3085 let twelve = build.iconst(i64, 12);
3086 let field = build.binary(Opcode::Add, args[0], twelve, Flags::default());
3087 build.load(Type::int(64), field, plain(), Flags::default());
3088
3089 // Two IR instructions and one machine instruction, which is what every read of a field
3090 // of a structure comes to.
3091 assert_eq!(
3092 lower(&mut names, &func),
3093 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
3094 %1:gpr = x64.mov_rm_64 [%0 + 12]\n}\n"
3095 );
3096 }
3097
3098 #[test]
3099 fn a_displacement_too_wide_to_encode_leaves_the_addition_where_it_is() {
3100 let i64 = Type::int(64);
3101 let (mut names, mut func, block, args) = blank(&[i64]);
3102 let mut build = Builder::new(&mut func, block);
3103 let big = build.iconst(i64, i128::from(i32::MAX) + 1);
3104 let far = build.binary(Opcode::Add, args[0], big, Flags::default());
3105 build.load(Type::int(32), far, plain(), Flags::default());
3106
3107 // A displacement is signed and 32 bits. The rule that folds one has a guard that turns
3108 // this down, so the addition stays and the load reads through what it produced. Nobody
3109 // wrote that fallback: it is the next way of showing the operand.
3110 let text = lower(&mut names, &func);
3111 assert!(text.contains("x64.mov_rm_32 [%2]"), "{text}");
3112 assert!(text.contains("x64.add_rr_64"), "{text}");
3113 }
3114
3115 #[test]
3116 fn a_store_of_a_value_that_was_loaded_is_two_instructions_and_no_arithmetic() {
3117 let i64 = Type::int(64);
3118 let (mut names, mut func, block, args) = blank(&[i64, i64]);
3119 let mut build = Builder::new(&mut func, block);
3120 let got = build.load(Type::int(8), args[0], plain(), Flags::default());
3121 build.store(got, args[1], plain(), Flags::default());
3122
3123 // A load feeding a store is the one place folding would be wrong: an x86-64 `mov` has at
3124 // most one memory operand, and there is no rule that takes two, so the load is left where
3125 // it is and the store reads the register it wrote.
3126 assert_eq!(
3127 lower(&mut names, &func),
3128 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
3129 %1:gpr($rsi) = x64.arg_val_64\n %2:gpr = x64.mov_rm_8 [%0]\n \
3130 x64.mov_mr_8 %2, [%1]\n}\n"
3131 );
3132 }
3133
3134 #[test]
3135 fn an_access_at_a_width_no_rule_is_written_at_is_reported() {
3136 let i64 = Type::int(64);
3137 let (mut names, mut source, block, args) = blank(&[i64]);
3138 let mut build = Builder::new(&mut source, block);
3139 build.load(Type::int(128), args[0], plain(), Flags::default());
3140
3141 // The width is the whole of what is wrong here, so the width is in the message: `load`
3142 // on its own is written about at every other width and would send a reader looking in
3143 // the wrong place.
3144 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
3145 .expect_err("nothing loads 128 bits");
3146 assert_eq!(failed.to_string(), "no rule lowers a `load` producing a `i128`");
3147 }
3148
3149 #[test]
3150 fn a_return_asks_for_the_value_in_the_register_the_caller_reads() {
3151 let (mut names, mut func, block, args) = blank(&[Type::int(32)]);
3152 let mut build = Builder::new(&mut func, block);
3153 build.ret(&[args[0]]);
3154
3155 // The register is not in the rule, the same way `cl` is not in the rule for a shift. It
3156 // is what the target says the instruction does with its operand, and the allocator is
3157 // what will act on it. There is no `ret` here, because giving the frame back has to
3158 // happen between this and leaving and the frame is not worked out yet.
3159 assert_eq!(
3160 lower(&mut names, &func),
3161 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
3162 x64.ret_val_32 %0($rax)\n}\n"
3163 );
3164 }
3165
3166 #[test]
3167 fn a_return_of_two_values_asks_for_the_second_register_as_well() {
3168 let i64 = Type::int(64);
3169 let (mut names, mut func, block, args) = blank(&[i64, i64]);
3170 let mut build = Builder::new(&mut func, block);
3171 build.ret(&[args[0], args[1]]);
3172
3173 // `struct { long a, b; } f(long a, long b)`, after the front end has classified it. Both
3174 // halves are integers, so the second is in the second integer return register, and both
3175 // pseudos say so the same way the one for a single value does.
3176 assert_eq!(
3177 lower(&mut names, &func),
3178 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
3179 %1:gpr($rsi) = x64.arg_val_64\n x64.ret_val_64 %0($rax)\n \
3180 x64.ret_val2_64 %1($rdx)\n}\n"
3181 );
3182 }
3183
3184 #[test]
3185 fn two_values_back_in_different_files_are_both_the_first_of_their_own() {
3186 let f64 = Type::float(rucc_ir::Float::F64);
3187 let (mut names, mut func, block, args) = blank(&[f64, Type::int(64)]);
3188 let mut build = Builder::new(&mut func, block);
3189 build.ret(&[args[0], args[1]]);
3190
3191 // `struct { double a; long b; } f(double a, long b)`. The two files are counted apart, so
3192 // neither half is the second of anything and the `double` is in `xmm0` rather than in the
3193 // register a second `double` would have been in. Getting this wrong is not a crash: the
3194 // caller reads a register nobody wrote, and this is where that is ruled out.
3195 assert_eq!(
3196 lower(&mut names, &func),
3197 "mfunc @f {\nblock0:\n %0:xmm($xmm0) = x64.arg_val_f64\n \
3198 %1:gpr($rdi) = x64.arg_val_64\n x64.ret_val_f64 %0($xmm0)\n \
3199 x64.ret_val_64 %1($rax)\n}\n"
3200 );
3201 }
3202
3203 #[test]
3204 fn two_of_the_same_file_back_take_the_first_two_of_it() {
3205 let f64 = Type::float(rucc_ir::Float::F64);
3206 let (mut names, mut func, block, args) = blank(&[f64, f64]);
3207 let mut build = Builder::new(&mut func, block);
3208 build.ret(&[args[0], args[1]]);
3209
3210 // `struct { double x, y; } f(double x, double y)`, which is the vector half of the pair
3211 // above and counts in its own file the same way.
3212 assert_eq!(
3213 lower(&mut names, &func),
3214 "mfunc @f {\nblock0:\n %0:xmm($xmm0) = x64.arg_val_f64\n \
3215 %1:xmm($xmm1) = x64.arg_val_f64\n x64.ret_val_f64 %0($xmm0)\n \
3216 x64.ret_val2_f64 %1($xmm1)\n}\n"
3217 );
3218 }
3219
3220 /// A function whose answer goes back through memory, with the pointer to the space for it in
3221 /// front of whatever else it takes. Only the signature says it is one.
3222 fn returning_through_memory(params: &[Type]) -> (Interner, Func, Block, Vec<Value>) {
3223 let mut names = Interner::new();
3224 let sret = Abi::Sret { size: 32, align: 8 };
3225 let mut signature = Signature::new().and_param(Param::with_abi(Type::PTR, sret));
3226 signature.params.extend(params.iter().copied().map(Param::new));
3227 let mut func = Func::new(names.intern("f"), signature);
3228 let block = func.create_block();
3229 let space = func.append_param(block, Type::PTR);
3230 let values = std::iter::once(space)
3231 .chain(params.iter().map(|&ty| func.append_param(block, ty)))
3232 .collect();
3233 (names, func, block, values)
3234 }
3235
3236 #[test]
3237 fn the_space_a_return_through_memory_was_given_goes_back_in_the_first_return_register() {
3238 let (mut names, mut func, block, _) = returning_through_memory(&[]);
3239 Builder::new(&mut func, block).ret(&[]);
3240
3241 // `struct big f(void)`, where `big` is too large to come back in registers. The `return`
3242 // carries nothing, because the value went into the space the caller handed over, and the
3243 // document still says that address comes back in `rax`. Nothing in the IR says it, so the
3244 // convention says it, and the pseudo is the one any other pointer return would use.
3245 assert_eq!(
3246 lower(&mut names, &func),
3247 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
3248 x64.ret_val_64 %0($rax)\n}\n"
3249 );
3250 }
3251
3252 #[test]
3253 fn what_the_function_did_in_between_does_not_take_the_register_off_it() {
3254 let (mut names, mut func, block, args) = returning_through_memory(&[Type::int(32)]);
3255 let mut build = Builder::new(&mut func, block);
3256 build.store(args[1], args[0], plain(), Flags::default());
3257 build.ret(&[]);
3258
3259 // The register is a read at the end and not a move at the start, so it is live across
3260 // everything between the two and the allocator has to keep it somewhere. In a function
3261 // with a call in it that somewhere is a callee saved register, and the address comes back
3262 // into `rax` here rather than whatever the last instruction happened to leave there. That
3263 // is issue #333, and a store is enough to show the value outlives the entry block.
3264 let text = lower(&mut names, &func);
3265 assert!(text.contains("x64.mov_mr_32 %1, [%0]"), "{text}");
3266 assert!(text.ends_with(" x64.ret_val_64 %0($rax)\n}\n"), "{text}");
3267 }
3268
3269 #[test]
3270 fn a_pointer_that_is_only_a_pointer_is_not_given_back() {
3271 let (mut names, mut func, block, args) = blank(&[Type::PTR]);
3272 let mut build = Builder::new(&mut func, block);
3273 build.store(args[0], args[0], plain(), Flags::default());
3274 build.ret(&[]);
3275
3276 // `void f(void **p)`. It takes a pointer first and returns nothing, which is the shape of
3277 // the one above and none of its meaning, and what tells them apart is the signature. A
3278 // `void` function leaves `rax` alone.
3279 assert!(!lower(&mut names, &func).contains("ret_val"));
3280 }
3281
3282 #[test]
3283 fn a_return_of_a_constant_puts_it_in_a_register_first() {
3284 let (mut names, mut func, block, _) = blank(&[]);
3285 let mut build = Builder::new(&mut func, block);
3286 let zero = build.iconst(Type::int(32), 0);
3287 build.ret(&[zero]);
3288
3289 // No rule returns an immediate, so the plan that offers one is turned down and the next
3290 // one materializes it. That is `int main(void) { return 0; }` in full, once the epilogue
3291 // is appended to it.
3292 assert_eq!(
3293 lower(&mut names, &func),
3294 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_ri_32 0\n x64.ret_val_32 %0($rax)\n}\n"
3295 );
3296 }
3297
3298 #[test]
3299 fn the_rule_that_writes_a_constant_down_is_recorded_as_a_rule_that_fired() {
3300 let (mut names, mut func, block, _) = blank(&[]);
3301 let mut build = Builder::new(&mut func, block);
3302 let zero = build.iconst(Type::int(32), 0);
3303 build.ret(&[zero]);
3304
3305 // The loop over the instructions passes a constant by, because a constant is written where
3306 // a register for it is first wanted rather than where the IR put it. So the only place a
3307 // rule about one is ever selected is the materialization, and a mark made in the loop
3308 // alone would report every rule about a constant as a rule nothing reaches.
3309 let out = super::func(&func, &mut names, &SYSV, &Elsewhere::default())
3310 .expect("every instruction has a rule");
3311 let rules = &crate::select::x86_64::TABLE.rules;
3312 let fired: Vec<&str> = rules
3313 .iter()
3314 .enumerate()
3315 .filter(|(index, _)| out.fired.has(*index))
3316 .map(|(_, rule)| rule.pattern)
3317 .collect();
3318 assert!(fired.contains(&"(iconst.i32 k)"), "{fired:?}");
3319 }
3320
3321 #[test]
3322 fn a_return_of_nothing_is_no_instruction_at_all() {
3323 let (mut names, mut func, block, _) = blank(&[]);
3324 let mut build = Builder::new(&mut func, block);
3325 build.ret(&[]);
3326
3327 // Every part of leaving a function that returns nothing is the epilogue's, and the
3328 // epilogue goes in after allocation. A block with nothing in it is the right answer here
3329 // rather than a function that could not be lowered.
3330 assert_eq!(lower(&mut names, &func), "mfunc @f {\nblock0:\n}\n");
3331 }
3332
3333 #[test]
3334 fn the_allocator_is_what_moves_the_answer_into_the_return_register() {
3335 let (mut names, mut source, block, _) = blank(&[]);
3336 let mut build = Builder::new(&mut source, block);
3337 let zero = build.iconst(Type::int(32), 0);
3338 build.ret(&[zero]);
3339
3340 let mut out = func(&source, &mut names, &SYSV, &Elsewhere::default())
3341 .expect("every instruction has a rule")
3342 .func;
3343 let env = env();
3344 let allocation = rucc_regalloc::run(&mut out, &env, "test");
3345 let frame = Frame::of(&out, &allocation, &Layout::new(&SYSV, REGS));
3346 finish(
3347 &mut out,
3348 &allocation,
3349 &frame,
3350 &Stack::default(),
3351 Convention::new(&SYSV, &FRAME),
3352 &mut names,
3353 );
3354
3355 // `int main(void) { return 0; }` end to end. Nothing here asked for `rax`: the rule said
3356 // the value goes back, the target said where, and the allocator is what made it true. The
3357 // epilogue is what leaves, and this function needs no frame, so it is the return alone.
3358 //
3359 // Two instructions and no copy, which is what a hint buys. The return insists on `rax`,
3360 // so `rax` is the register the allocator tries first for the value the return reads, and
3361 // the constant is written straight into it.
3362 assert_eq!(
3363 mir::print_func(&out, &names, ®S),
3364 "mfunc @f {\nblock0:\n $rax = x64.mov_ri_32 0\n \
3365 x64.ret_val_32 $rax($rax)\n x64.ret\n}\n"
3366 );
3367 }
3368
3369 #[test]
3370 fn a_function_of_two_arguments_is_a_whole_function_now() {
3371 let i32 = Type::int(32);
3372 let (mut names, mut source, block, args) = blank(&[i32, i32]);
3373 let mut build = Builder::new(&mut source, block);
3374 let sum = build.binary(Opcode::Add, args[0], args[1], Flags::default());
3375 build.ret(&[sum]);
3376
3377 let mut out = func(&source, &mut names, &SYSV, &Elsewhere::default())
3378 .expect("every instruction has a rule")
3379 .func;
3380 let env = env();
3381 let allocation = rucc_regalloc::run(&mut out, &env, "test");
3382 let frame = Frame::of(&out, &allocation, &Layout::new(&SYSV, REGS));
3383 finish(
3384 &mut out,
3385 &allocation,
3386 &frame,
3387 &Stack::default(),
3388 Convention::new(&SYSV, &FRAME),
3389 &mut names,
3390 );
3391
3392 // `int f(int a, int b) { return a + b; }` end to end, and this is the test the argument
3393 // side exists for. Before it there was no way to write one: the allocator refuses a
3394 // function whose entry block takes parameters, because there is no edge into an entry
3395 // block for the moves that give a block parameter its value to go on.
3396 //
3397 // One move, and it is the one the machine's addition needs rather than one the allocator
3398 // owes anybody. Each argument stays in the register it arrived in, because the pseudo
3399 // that defines it insists on that register and the allocator now tries it first, and the
3400 // sum stays in the register the addition wrote it to until the return reads it out. The
3401 // copy in front of a two address instruction is what makes its destination one of the
3402 // registers it reads, and the source operand keeps its own name because the destination
3403 // is what the encoder writes.
3404 assert_eq!(
3405 mir::print_func(&out, &names, ®S),
3406 "mfunc @f {\nblock0:\n $rdi($rdi) = x64.arg_val_32\n \
3407 $rsi($rsi) = x64.arg_val_32\n \
3408 $rdi(reuse 1) = x64.add_rr_32 $rdi, $rsi\n $rax = x64.mov_rr_64 $rdi\n \
3409 x64.ret_val_32 $rax($rax)\n x64.ret\n}\n"
3410 );
3411 }
3412
3413 #[test]
3414 fn an_argument_with_no_register_left_for_it_is_read_out_of_the_caller_s_stack() {
3415 let i64 = Type::int(64);
3416 let (mut names, mut source, block, args) = blank(&[i64; 7]);
3417 let mut build = Builder::new(&mut source, block);
3418 build.ret(&[args[6]]);
3419
3420 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
3421 .expect("the seventh is read from memory");
3422
3423 // SysV passes six integers in registers and the seventh in the caller's memory, so six of
3424 // these are pseudos that encode to nothing and the seventh is a load that encodes to real
3425 // bytes. Its displacement is nothing here for the reason a local's is: there is no frame
3426 // yet. What the walk hands on is which instruction is waiting, and for how far up the
3427 // caller's argument area, which is the bottom of it because it is the first one there.
3428 assert_eq!(lowered.stack.arguments.len(), 1);
3429 assert_eq!(lowered.stack.arguments[0].1, 0);
3430 let text = mir::print_func(&lowered.func, &names, ®S);
3431 assert!(text.contains("%6:gpr = x64.mov_rm_64 [$rsp]"), "{text}");
3432 assert_eq!(text.matches("x64.arg_val_64").count(), 6, "{text}");
3433 }
3434
3435 #[test]
3436 fn the_frame_is_what_says_how_far_up_the_caller_s_stack_an_argument_is() {
3437 let i64 = Type::int(64);
3438 let (mut names, mut source, block, args) = blank(&[i64; 8]);
3439 let mut build = Builder::new(&mut source, block);
3440 let sum = build.binary(Opcode::Add, args[6], args[7], Flags::default());
3441 build.ret(&[sum]);
3442
3443 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
3444 .expect("both are read from memory");
3445 let stack = lowered.stack;
3446 let mut out = lowered.func;
3447 let env = env();
3448 let allocation = rucc_regalloc::run(&mut out, &env, "test");
3449 let layout = stack.layout(Layout::new(&SYSV, REGS));
3450 let frame = Frame::of(&out, &allocation, &layout);
3451 finish(&mut out, &allocation, &frame, &stack, Convention::new(&SYSV, &FRAME), &mut names);
3452
3453 // A leaf that takes no frame, so the stack pointer never moves and the only thing between
3454 // it and the caller's arguments is the return address the call pushed. The seventh
3455 // parameter is at the bottom of the caller's argument area and the eighth is one word
3456 // further up, which is the eight bytes between the two offsets.
3457 let text = mir::print_func(&out, &names, ®S);
3458 assert_eq!(frame.size(), 0);
3459 assert_eq!(frame.incoming(), Incoming::from_stack(8));
3460 assert!(text.contains("x64.mov_rm_64 [$rsp + 8]"), "{text}");
3461 assert!(text.contains("x64.mov_rm_64 [$rsp + 16]"), "{text}");
3462 }
3463
3464 #[test]
3465 fn a_realigned_frame_reaches_the_caller_s_arguments_through_the_frame_pointer() {
3466 let i64 = Type::int(64);
3467 let (mut names, mut source, block, args) = blank(&[i64; 7]);
3468 let wide = slot(&mut source, block, 64, 32);
3469 let mut build = Builder::new(&mut source, block);
3470 build.store(args[6], wide, plain(), Flags::default());
3471 build.ret(&[args[6]]);
3472
3473 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
3474 .expect("every instruction has a rule");
3475 let stack = lowered.stack;
3476 let mut out = lowered.func;
3477 let env = env();
3478 let allocation = rucc_regalloc::run(&mut out, &env, "test");
3479 let layout = stack.layout(Layout::new(&SYSV, REGS));
3480 let frame = Frame::of(&out, &allocation, &layout);
3481 finish(&mut out, &allocation, &frame, &stack, Convention::new(&SYSV, &FRAME), &mut names);
3482
3483 // A local wanting thirty two byte alignment makes the prologue force the stack pointer,
3484 // which throws away how far the caller's stack was. So the load the lowering wrote off the
3485 // stack pointer is rewritten to read through the frame pointer, at the one distance that
3486 // survives: the word the prologue pushed the frame pointer into, and the return address
3487 // above it.
3488 let text = mir::print_func(&out, &names, ®S);
3489 assert_eq!(frame.realign(), Some(32));
3490 assert_eq!(frame.incoming(), Incoming::from_frame(16));
3491 assert!(text.contains("x64.mov_rm_64 [$rbp + 16]"), "{text}");
3492 assert!(!text.contains("x64.mov_rm_64 [$rsp"), "{text}");
3493 }
3494
3495 #[test]
3496 fn a_jump_is_the_edge_and_nothing_else() {
3497 let i32 = Type::int(32);
3498 let (mut names, mut source, entry, args) = blank(&[i32]);
3499 let next = source.create_block();
3500 let got = source.append_param(next, i32);
3501 Builder::new(&mut source, entry).jump(next, &[args[0]]);
3502 Builder::new(&mut source, next).ret(&[got]);
3503
3504 // Two blocks and two instructions, and the jump is neither of them. What it was is the
3505 // arm on the first block, and what the arm carries is the argument it was called with.
3506 assert_eq!(
3507 lower(&mut names, &source),
3508 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32 block1(%0)\n\n\
3509 block1(%1:gpr):\n x64.ret_val_32 %1($rax)\n}\n"
3510 );
3511 }
3512
3513 /// A block that reads what a block below it writes is filled after it, not before it.
3514 ///
3515 /// The blocks are written entry, `early`, `late`, `exit`, and the entry jumps straight past
3516 /// `early` to `late`, so `late` dominates `early` while sitting below it in the function.
3517 /// Filling them in the order they are written reaches the read in `early` first, and reading
3518 /// a value with no register yet mints one. The cast in `late` is no instruction at all, so
3519 /// what it does is give its answer the register its operand is already in, and that is not
3520 /// the register the read minted. Nothing writes the register the read minted. The printer
3521 /// says `%?` for a register nothing defines, which is what this looks for, and what came out
3522 /// of the real bug was SQLite loading a stack slot no store ever reached.
3523 #[test]
3524 fn a_block_that_reads_what_a_block_below_it_writes_is_filled_after_it() {
3525 let i64 = Type::int(64);
3526 let (mut names, mut source, entry, args) = blank(&[i64, i64]);
3527 let early = source.create_block();
3528 let late = source.create_block();
3529 let exit = source.create_block();
3530
3531 Builder::new(&mut source, entry).jump(late, &[]);
3532 let ptr = cast(&mut source, late, Opcode::IntToPtr, args[0], Type::PTR);
3533 Builder::new(&mut source, early).ret(&[ptr]);
3534 let mut build = Builder::new(&mut source, late);
3535 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
3536 build.br_if(cond, early, &[], exit, &[]);
3537 Builder::new(&mut source, exit).ret(&[args[1]]);
3538
3539 let text = lower(&mut names, &source);
3540 assert!(!text.contains("%?"), "every register has something that writes it: {text}");
3541 }
3542
3543 /// A constant is written where it is wanted rather than where the IR defined it, and two
3544 /// blocks wanting the same one is two places. Writing it once and reading it in both is a
3545 /// register read where nothing wrote it, unless the block it was written in happens to
3546 /// dominate the other, which nothing here checks and which the second arm of a branch never
3547 /// does. Each block gets its own copy of the number instead.
3548 #[test]
3549 fn a_constant_two_blocks_want_is_written_in_both_of_them() {
3550 let i32 = Type::int(32);
3551 let (mut names, mut source, entry, args) = blank(&[i32, i32]);
3552 let then = source.create_block();
3553 let other = source.create_block();
3554 let join = source.create_block();
3555 let got = source.append_param(join, i32);
3556
3557 let mut build = Builder::new(&mut source, entry);
3558 let seven = build.iconst(i32, 7);
3559 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
3560 build.br_if(cond, then, &[], other, &[]);
3561 // Both arms want the seven in a register, because a block argument is never an immediate,
3562 // and neither arm dominates the other.
3563 Builder::new(&mut source, then).jump(join, &[seven]);
3564 Builder::new(&mut source, other).jump(join, &[seven]);
3565 Builder::new(&mut source, join).ret(&[got]);
3566
3567 let text = lower(&mut names, &source);
3568 assert_eq!(text.matches("x64.mov_ri_32 7").count(), 2, "one seven per block: {text}");
3569 }
3570
3571 /// An argument on an edge out of a block that leaves two ways is read after every instruction
3572 /// of the block is written, and reading one can write an instruction, which would land after
3573 /// the branch that has already jumped past it. The branch goes back on the end.
3574 #[test]
3575 fn a_constant_an_edge_wants_is_written_before_the_branch_and_not_after_it() {
3576 let i32 = Type::int(32);
3577 let (mut names, mut source, entry, args) = blank(&[i32, i32]);
3578 let then = source.create_block();
3579 let join = source.create_block();
3580 let got = source.append_param(join, i32);
3581
3582 let mut build = Builder::new(&mut source, entry);
3583 let nine = build.iconst(i32, 9);
3584 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
3585 build.br_if(cond, then, &[], join, &[nine]);
3586 Builder::new(&mut source, then).jump(join, &[args[0]]);
3587 Builder::new(&mut source, join).ret(&[got]);
3588
3589 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
3590 .expect("every instruction has a rule")
3591 .func;
3592 let entry = out.entry().expect("an entry block");
3593 let last = out.terminator(entry).expect("a block that leaves two ways has a branch");
3594 let branch = names.intern("x64.br_cond_8");
3595 assert_eq!(
3596 out[last].opcode,
3597 mir::Opcode::new(branch),
3598 "the branch is last: {}",
3599 mir::print_func(&out, &names, ®S)
3600 );
3601 }
3602
3603 #[test]
3604 fn a_conditional_branch_is_lowered_to_the_condition_and_nothing_about_where_it_goes() {
3605 let i32 = Type::int(32);
3606 let (mut names, mut source, entry, args) = blank(&[i32, i32]);
3607 let then = source.create_block();
3608 let other = source.create_block();
3609 let mut build = Builder::new(&mut source, entry);
3610 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
3611 build.br_if(cond, then, &[], other, &[]);
3612 Builder::new(&mut source, then).ret(&[args[0]]);
3613 Builder::new(&mut source, other).ret(&[args[1]]);
3614
3615 // The comparison writes a byte and the branch reads it, and neither says a block. Both
3616 // arms are on the entry block, in the order the branch took them, so the arm that runs
3617 // when the condition holds is the first.
3618 assert_eq!(
3619 lower(&mut names, &source),
3620 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
3621 %1:gpr($rsi) = x64.arg_val_32\n %2:gpr = x64.cmp_set_l_32 %0, %1\n \
3622 x64.br_cond_8 %2, block1, block2\n\n\
3623 block1:\n x64.ret_val_32 %0($rax)\n\n\
3624 block2:\n x64.ret_val_32 %1($rax)\n}\n"
3625 );
3626 }
3627
3628 /// A choice between two values, which is one instruction and no blocks at all.
3629 ///
3630 /// The arms come out the other way round from the IR, because a conditional move overwrites its
3631 /// destination and the destination is the arm taken when the condition does not hold. The
3632 /// condition arrives last for the same reason: it is read by the test in front of the move
3633 /// rather than by the move.
3634 #[test]
3635 fn a_select_is_lowered_to_a_test_and_a_conditional_move() {
3636 let i32 = Type::int(32);
3637 let (mut names, mut source, entry, args) = blank(&[i32, i32]);
3638 let mut build = Builder::new(&mut source, entry);
3639 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
3640 let picked = build.select(cond, args[0], args[1]);
3641 build.ret(&[picked]);
3642
3643 assert_eq!(
3644 lower(&mut names, &source),
3645 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
3646 %1:gpr($rsi) = x64.arg_val_32\n %2:gpr = x64.cmp_set_l_32 %0, %1\n \
3647 %3:gpr(reuse 1) = x64.test_cmov_ne_32 %1, %0, %2\n \
3648 x64.ret_val_32 %3($rax)\n}\n"
3649 );
3650 }
3651
3652 #[test]
3653 fn a_branch_over_a_block_is_a_whole_function_now() {
3654 let i32 = Type::int(32);
3655 let (mut names, mut source, entry, args) = blank(&[i32, i32]);
3656 let then = source.create_block();
3657 let other = source.create_block();
3658 let join = source.create_block();
3659 let got = source.append_param(join, i32);
3660 let mut build = Builder::new(&mut source, entry);
3661 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
3662 build.br_if(cond, then, &[], other, &[]);
3663 let mut build = Builder::new(&mut source, then);
3664 let sum = build.binary(Opcode::Add, args[0], args[1], Flags::default());
3665 build.jump(join, &[sum]);
3666 Builder::new(&mut source, other).jump(join, &[args[1]]);
3667 Builder::new(&mut source, join).ret(&[got]);
3668
3669 // `int f(int a, int b) { if (a < b) return a + b; else return b; }` end to end, written
3670 // the way a front end writes it: both arms of the branch are blocks of their own and the
3671 // return is the block they meet at. No edge here is critical, because the two arms out of
3672 // the entry carry nothing and the two arms into the join each leave a block that goes
3673 // nowhere else, so each has its own end to put its move at.
3674 let mut out = func(&source, &mut names, &SYSV, &Elsewhere::default())
3675 .expect("every instruction has a rule")
3676 .func;
3677 assert_eq!(crate::split::critical(&mut out), 0, "no edge here is critical");
3678 let env = env();
3679 let allocation = rucc_regalloc::run(&mut out, &env, "test");
3680 let frame = Frame::of(&out, &allocation, &Layout::new(&SYSV, REGS));
3681 finish(
3682 &mut out,
3683 &allocation,
3684 &frame,
3685 &Stack::default(),
3686 Convention::new(&SYSV, &FRAME),
3687 &mut names,
3688 );
3689
3690 // One epilogue, on the join, which is the one block the function leaves from, and the
3691 // moves that give the join its parameter are at the end of each arm. Every register is
3692 // physical and the branch is still a branch on a register, because turning it into a
3693 // `test` and a `jcc` is the block layout's and there is no block layout yet.
3694 let text = mir::print_func(&out, &names, ®S);
3695 assert_eq!(text.matches("x64.ret\n").count(), 1, "{text}");
3696 assert!(text.contains("x64.br_cond_8"), "{text}");
3697 assert!(text.contains("x64.add_rr_32"), "{text}");
3698 assert!(!text.contains('%'), "{text}");
3699 }
3700
3701 #[test]
3702 fn a_critical_edge_is_split_before_the_allocator_ever_sees_it() {
3703 let i32 = Type::int(32);
3704 let (mut names, mut source, entry, args) = blank(&[i32, i32]);
3705 let then = source.create_block();
3706 let join = source.create_block();
3707 let got = source.append_param(join, i32);
3708 let mut build = Builder::new(&mut source, entry);
3709 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
3710 build.br_if(cond, then, &[], join, &[args[1]]);
3711 Builder::new(&mut source, then).jump(join, &[args[0]]);
3712 let mut build = Builder::new(&mut source, join);
3713 let twice = build.binary(Opcode::Add, got, got, Flags::default());
3714 build.ret(&[twice]);
3715
3716 // The else arm is critical: the entry block leaves two ways and the join is arrived at
3717 // two ways, and the arm carries a value. Without splitting it the allocator asserts,
3718 // because the move that gives the join its parameter would have to run at the end of a
3719 // block that also goes to the other arm.
3720 let mut out = func(&source, &mut names, &SYSV, &Elsewhere::default())
3721 .expect("every instruction has a rule")
3722 .func;
3723 assert_eq!(crate::split::critical(&mut out), 1);
3724 let env = env();
3725 let allocation = rucc_regalloc::run(&mut out, &env, "test");
3726 let frame = Frame::of(&out, &allocation, &Layout::new(&SYSV, REGS));
3727 finish(
3728 &mut out,
3729 &allocation,
3730 &frame,
3731 &Stack::default(),
3732 Convention::new(&SYSV, &FRAME),
3733 &mut names,
3734 );
3735
3736 // The block the split added is where the move went, and it is the whole of that block.
3737 let text = mir::print_func(&out, &names, ®S);
3738 assert_eq!(out.block_count(), 4, "{text}");
3739 assert_eq!(text.matches("x64.ret\n").count(), 1, "{text}");
3740 }
3741
3742 #[test]
3743 fn a_call_passes_what_the_convention_says_and_takes_back_what_it_says() {
3744 let i32 = Type::int(32);
3745 let (mut names, mut source, block, args) = blank(&[i32, i32]);
3746 let sig =
3747 source.add_signature(Signature::new().with_params(&[i32, i32]).with_returns(&[i32]));
3748 let callee = names.intern("g");
3749 let call = Builder::new(&mut source, block).call(callee, sig, &[args[0], args[1]]);
3750 let got = source[call].first_result.expect("an integer comes back");
3751 Builder::new(&mut source, block).ret(&[got]);
3752
3753 // `int f(int a, int b) { return g(a, b); }`. The arguments arrived where the call wants
3754 // them, so what the call reads is what arrived, and the whole of the convention is in the
3755 // constraints rather than in a move.
3756 let text = lower(&mut names, &source);
3757 assert!(text.contains("= x64.call %0($rdi), %1($rsi), @g"), "{text}");
3758 assert!(text.contains("x64.ret_val_32 %2($rax)"), "{text}");
3759 // What the call writes is the value that comes back and then every register the callee is
3760 // free to destroy, in both classes, which is the whole of what stops the allocator from
3761 // leaving something in one of them.
3762 assert!(text.contains("%2:gpr($rax), $rcx, $rdx, $r8, $r9, $r10, $r11, $xmm0,"), "{text}");
3763 assert!(text.contains("$xmm15 = x64.call"), "{text}");
3764 }
3765
3766 #[test]
3767 fn what_the_frame_owes_a_call_comes_back_with_the_function() {
3768 let i32 = Type::int(32);
3769 let sig = |source: &mut Func| source.add_signature(Signature::new().with_params(&[i32]));
3770
3771 let (mut names, mut source, block, args) = blank(&[i32]);
3772 let sig = sig(&mut source);
3773 let callee = names.intern("g");
3774 Builder::new(&mut source, block).call(callee, sig, &[args[0]]);
3775 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
3776 .expect("every instruction has a rule");
3777
3778 // Nothing on the stack, so nothing owed, but not a leaf either: a function that calls
3779 // owes the callee an aligned stack pointer and may not use the red zone.
3780 assert_eq!(out.stack.calls, Some(0));
3781 let layout = out.stack.layout(Layout::new(&SYSV, REGS));
3782 assert!(!layout.leaf);
3783 assert_eq!(layout.outgoing, 0);
3784
3785 // The same call under the other convention owes thirty two bytes for the callee to spill
3786 // its register arguments into, which is a fact about the convention and not about the call.
3787 let out = func(&source, &mut names, &x86_64::WIN64, &Elsewhere::default())
3788 .expect("every instruction has a rule");
3789 assert_eq!(out.stack.calls, Some(32));
3790
3791 // And a function that calls nothing is a leaf, which is what says it may use the red zone.
3792 let (mut names, mut source, block, args) = blank(&[i32]);
3793 Builder::new(&mut source, block).ret(&[args[0]]);
3794 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
3795 .expect("every instruction has a rule");
3796 assert_eq!(out.stack.calls, None);
3797 assert!(out.stack.layout(Layout::new(&SYSV, REGS)).leaf);
3798 }
3799
3800 #[test]
3801 fn a_value_that_outlives_a_call_is_not_left_where_the_call_destroys_it() {
3802 let i32 = Type::int(32);
3803 let (mut names, mut source, block, args) = blank(&[i32]);
3804 let sig = source.add_signature(Signature::new().with_params(&[i32]).with_returns(&[i32]));
3805 let callee = names.intern("g");
3806 let call = Builder::new(&mut source, block).call(callee, sig, &[args[0]]);
3807 let got = source[call].first_result.expect("an integer comes back");
3808 let mut build = Builder::new(&mut source, block);
3809 let sum = build.binary(Opcode::Add, got, args[0], Flags::default());
3810 build.ret(&[sum]);
3811
3812 // `int f(int a) { return g(a) + a; }`, which is the smallest program that asks the
3813 // question: `a` is read after the call and `rdi` is a register the call destroys.
3814 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
3815 .expect("every instruction has a rule");
3816 let layout = lowered.stack.layout(Layout::new(&SYSV, REGS));
3817 let mut out = lowered.func;
3818 let env = env();
3819 let allocation = rucc_regalloc::run(&mut out, &env, "test");
3820 let frame = Frame::of(&out, &allocation, &layout);
3821 finish(
3822 &mut out,
3823 &allocation,
3824 &frame,
3825 &Stack::default(),
3826 Convention::new(&SYSV, &FRAME),
3827 &mut names,
3828 );
3829
3830 // It went to a register the callee has to put back, and the prologue and epilogue are what
3831 // put it back, which is the whole bargain the two halves of a convention make.
3832 let text = mir::print_func(&out, &names, ®S);
3833 assert!(text.contains("$rbx"), "{text}");
3834 assert!(!text.contains('%'), "{text}");
3835 assert_eq!(text.matches("x64.call").count(), 1, "{text}");
3836 }
3837
3838 #[test]
3839 fn a_call_with_more_arguments_than_registers_writes_the_rest_into_the_outgoing_area() {
3840 let i64 = Type::int(64);
3841 let (mut names, mut source, block, args) = blank(&[i64]);
3842 let seven = vec![i64; 7];
3843 let sig = source.add_signature(Signature::new().with_params(&seven));
3844 let callee = names.intern("g");
3845 let passed = vec![args[0]; 7];
3846 Builder::new(&mut source, block).call(callee, sig, &passed);
3847
3848 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
3849 .expect("the seventh goes to memory");
3850 // The bytes the call needs are on the layout the frame is worked out from, so that the
3851 // frame reserves as many as the widest call in the function asked for.
3852 assert_eq!(lowered.stack.calls, Some(8));
3853 let text = mir::print_func(&lowered.func, &names, ®S);
3854 assert!(text.contains("x64.mov_mr_64 %0, [$rsp]\n"), "{text}");
3855 }
3856
3857 #[test]
3858 fn a_call_this_cannot_make_is_reported_rather_than_made() {
3859 let (mut names, mut source, block, _) = blank(&[]);
3860 let returns = [Type::float(rucc_ir::Float::F80), Type::int(64)];
3861 let sig = source.add_signature(Signature::new().with_returns(&returns));
3862 let callee = names.intern("g");
3863 Builder::new(&mut source, block).call(callee, sig, &[]);
3864 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
3865 .expect_err("a long double is on the x87");
3866 assert_eq!(failed.to_string(), "what this call gives back is on the x87 stack");
3867 }
3868
3869 /// A `long double` on its own is a different answer, because on its own it comes back on the
3870 /// x87 stack rather than in a register, which is somewhere the call cannot be said to write.
3871 ///
3872 /// So the call gives back nothing at all and the value is taken off the stack by the `fstp`
3873 /// straight after it. That instruction has to be straight after it: the stack is one place and
3874 /// anything else that touched it before this ran would be looking at the value still on it.
3875 #[test]
3876 fn a_call_that_gives_back_a_long_double_takes_it_off_the_stack_at_once() {
3877 let (mut names, mut source, block, _) = blank(&[]);
3878 let long_double = Type::float(rucc_ir::Float::F80);
3879 let sig = source.add_signature(Signature::new().with_returns(&[long_double]));
3880 let callee = names.intern("g");
3881 Builder::new(&mut source, block).call(callee, sig, &[]);
3882
3883 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
3884 .expect("the value comes back in st0");
3885 let text = mir::print_func(&lowered.func, &names, ®S);
3886 let after: Vec<&str> =
3887 text.lines().skip_while(|line| !line.contains("x64.call")).skip(1).collect();
3888 assert_eq!(after[0].trim(), "%0:gpr = x64.lea_64 [$rsp]", "{text}");
3889 assert_eq!(after[1].trim(), "x64.fstp_t [%0]", "{text}");
3890 // And the slot it went into is the sixteen bytes the type takes, like every other one.
3891 assert_eq!(lowered.stack.locals.len(), 1, "{text}");
3892 assert_eq!(lowered.stack.locals[0].size, X87_BYTES);
3893 }
3894
3895 #[test]
3896 fn a_call_through_an_address_goes_through_the_register_the_address_is_in() {
3897 let i32 = Type::int(32);
3898 let (mut names, mut source, block, args) = blank(&[Type::PTR, i32]);
3899 let sig = source.add_signature(Signature::new().with_params(&[i32]).with_returns(&[i32]));
3900 let varargs = source.push_abis(&[]);
3901 let info = source.add_call(CallInfo { callee: None, signature: sig, varargs });
3902 let mut build = Builder::new(&mut source, block);
3903 let inst = InstData {
3904 args: build.func().push_values(&[args[0], args[1]]),
3905 extra: Extra::Call(info),
3906 ..InstData::new(Opcode::CallIndirect)
3907 };
3908 let called = build.inst(inst, &[i32]);
3909 let got = source[called].first_result.expect("an integer comes back");
3910 Builder::new(&mut source, block).ret(&[got]);
3911
3912 // `int f(int (*g)(int), int a) { return g(a); }`. The first operand is the address and
3913 // the arguments are the ones behind it, and everything else about the call is what a call
3914 // to a name would have been.
3915 let text = lower(&mut names, &source);
3916 assert!(text.contains("= x64.call_reg %0, %1($rdi)"), "{text}");
3917 assert!(text.contains("x64.ret_val_32 %2($rax)"), "{text}");
3918 assert!(!text.contains("@g"), "a call through an address names nobody: {text}");
3919 }
3920
3921 #[test]
3922 fn an_instruction_no_rule_covers_is_reported() {
3923 let (mut names, mut source, block, args) = blank(&[Type::PTR]);
3924 let mut build = Builder::new(&mut source, block);
3925 let operands = build.func().push_values(&[args[0]]);
3926 build.inst(InstData { args: operands, ..InstData::new(Opcode::LongjmpMarker) }, &[]);
3927
3928 // The mark that a jump goes back through here, which nothing writes an instruction for
3929 // yet: what it needs is for the allocator to be told a block can be arrived at twice, and
3930 // that is `tamnd/rucc#223`. Nothing about it is a width or a register, so there is nothing
3931 // for the message to add beyond the name.
3932 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
3933 .expect_err("no rule writes a longjmp marker");
3934 assert_eq!(failed.to_string(), "no rule lowers a `longjmp_marker`");
3935
3936 // It produces nothing, so there is no type in the message and nothing invents one, and the
3937 // instruction comes back so a caller can ask the function where it was.
3938 let inst = failed.inst().expect("the instruction it is about");
3939 assert_eq!(source[inst].opcode, Opcode::LongjmpMarker);
3940 }
3941
3942 /// A barrier is written by name here, and what it is depends on the ordering and on nothing
3943 /// else. `crate::expand` is where the reasoning about this machine's memory model lives.
3944 #[test]
3945 fn a_barrier_is_one_instruction_at_the_strongest_ordering_and_none_below_it() {
3946 for order in MemOrder::all().filter(|&order| order != MemOrder::NotAtomic) {
3947 let (mut names, mut source, block, _) = blank(&[]);
3948 let mut build = Builder::new(&mut source, block);
3949 build
3950 .inst(InstData { extra: Extra::Order(order), ..InstData::new(Opcode::Fence) }, &[]);
3951
3952 let text = lower(&mut names, &source);
3953 assert_eq!(text.contains("x64.mfence"), order == MemOrder::SeqCst, "{order:?}: {text}");
3954 }
3955 }
3956
3957 /// A compare and exchange is written by name too, and at the width of the value rather than at
3958 /// the width of the address, which is the mistake worth pinning: everything here is a pointer
3959 /// and only the value says how many bytes the instruction touches.
3960 #[test]
3961 fn a_compare_and_exchange_is_one_instruction_at_the_width_of_the_value() {
3962 for bits in [8, 16, 32, 64] {
3963 let ty = Type::int(bits);
3964 let (mut names, mut source, block, args) = blank(&[Type::PTR, ty, ty]);
3965 let mut build = Builder::new(&mut source, block);
3966 let mem = build.func().add_mem(MemInfo {
3967 size: u64::from(bits / 8),
3968 align: bits / 8,
3969 order: MemOrder::SeqCst,
3970 ..plain()
3971 });
3972 let operands = build.func().push_values(&[args[0], args[1], args[2]]);
3973 build.inst(
3974 InstData {
3975 args: operands,
3976 extra: Extra::Mem(mem),
3977 ..InstData::new(Opcode::Cmpxchg)
3978 },
3979 &[ty, Type::I1],
3980 );
3981
3982 // Two values out of one instruction, the first of them in the register the machine
3983 // reads the expected value out of, the second free for the allocator to place. The
3984 // address is the memory operand and neither of the two values is.
3985 let text = lower(&mut names, &source);
3986 let written = format!("%3:gpr($rax), %4:gpr = x64.cmpxchg_{bits} %1($rax), %2, [%0]");
3987 assert!(text.contains(&written), "{bits}: {text}");
3988 }
3989 }
3990
3991 #[test]
3992 fn more_values_back_than_the_convention_has_registers_for_is_reported() {
3993 let i64 = Type::int(64);
3994 let (mut names, mut source, block, args) = blank(&[i64, i64, i64]);
3995 let mut build = Builder::new(&mut source, block);
3996 build.ret(&[args[0], args[1], args[2]]);
3997
3998 // Two integers come back in `rax` and `rdx` and a third has nowhere to go, which is not a
3999 // gap in the rules but the convention saying no. The front end classifies before it gets
4000 // here, so this is the shape that would mean the classification went wrong.
4001 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4002 .expect_err("only two come back");
4003 assert_eq!(
4004 failed.to_string(),
4005 "what this function gives back takes more registers than this convention has for it"
4006 );
4007
4008 let inst = failed.inst().expect("the instruction it is about");
4009 assert_eq!(source[inst].opcode, Opcode::Return);
4010 }
4011
4012 /// A refusal about a signature has no instruction, which is what makes it the one arm apart.
4013 ///
4014 /// Everything else is about something written somewhere in the body and hands it back so a
4015 /// caller can ask the function where it came from. A parameter arrives before the first
4016 /// instruction runs, so there is nothing in the body to point at and the message is about
4017 /// the function.
4018 #[test]
4019 fn a_refusal_about_a_parameter_has_no_instruction_to_point_at() {
4020 let missing = Unsupported::Argument { index: 0, missing: Missing::OnX87 };
4021 assert_eq!(missing.inst(), None);
4022 }
4023
4024 /// An `alloca` of a fixed size, which is what every local whose address is taken becomes.
4025 fn slot(source: &mut Func, block: Block, size: u64, align: u32) -> Value {
4026 let info = MemInfo { size, align, ..plain() };
4027 let mut build = Builder::new(source, block);
4028 let mem = build.func().add_mem(info);
4029 build.value(InstData { extra: Extra::Mem(mem), ..InstData::new(Opcode::Alloca) }, Type::PTR)
4030 }
4031
4032 #[test]
4033 fn a_local_is_memory_in_the_frame_and_one_instruction_that_says_where() {
4034 let (mut names, mut source, block, _) = blank(&[]);
4035 let slot = slot(&mut source, block, 4, 4);
4036 let mut build = Builder::new(&mut source, block);
4037 let nine = build.iconst(Type::int(32), 9);
4038 build.store(nine, slot, plain(), Flags::default());
4039 let loaded = build.load(Type::int(32), slot, plain(), Flags::default());
4040 build.ret(&[loaded]);
4041
4042 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4043 .expect("every instruction has a rule");
4044
4045 // Four bytes on the list the frame is laid out from, and the one instruction that reads
4046 // where they went. Its displacement is nothing here because there is no frame yet, and
4047 // which instruction is waiting for which local is what `finish` is handed.
4048 assert_eq!(lowered.stack.locals, vec![Local { size: 4, align: 4 }]);
4049 assert_eq!(lowered.stack.addresses.len(), 1);
4050 assert_eq!(lowered.stack.addresses[0].1, 0);
4051 assert_eq!(
4052 mir::print_func(&lowered.func, &names, ®S),
4053 "mfunc @f {\nblock0:\n %0:gpr = x64.lea_64 [$rsp]\n \
4054 %1:gpr = x64.mov_ri_32 9\n x64.mov_mr_32 %1, [%0]\n \
4055 %2:gpr = x64.mov_rm_32 [%0]\n x64.ret_val_32 %2($rax)\n}\n"
4056 );
4057 }
4058
4059 #[test]
4060 fn the_frame_is_what_fills_the_address_of_a_local_in() {
4061 let (mut names, mut source, block, _) = blank(&[]);
4062 let slot = slot(&mut source, block, 4, 4);
4063 let mut build = Builder::new(&mut source, block);
4064 let nine = build.iconst(Type::int(32), 9);
4065 build.store(nine, slot, plain(), Flags::default());
4066 let loaded = build.load(Type::int(32), slot, plain(), Flags::default());
4067 build.ret(&[loaded]);
4068
4069 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4070 .expect("every instruction has a rule");
4071 let stack = lowered.stack;
4072 let mut out = lowered.func;
4073 let env = env();
4074 let allocation = rucc_regalloc::run(&mut out, &env, "test");
4075 let layout = stack.layout(Layout::new(&SYSV, REGS));
4076 let frame = Frame::of(&out, &allocation, &layout);
4077 finish(&mut out, &allocation, &frame, &stack, Convention::new(&SYSV, &FRAME), &mut names);
4078
4079 // `int f(void) { int x; x = 9; return x; }` with the address of `x` taken, end to end.
4080 // A leaf small enough to live in the red zone takes no frame at all, so the stack pointer
4081 // never moves and the four bytes are below it, which is what the negative offset is. The
4082 // instruction the lowering left with nothing in its displacement now has the answer in it.
4083 let text = mir::print_func(&out, &names, ®S);
4084 assert!(text.contains("$rax = x64.lea_64 [$rsp - 8]"), "{text}");
4085 assert!(!text.contains("x64.sub_ri_64"), "{text}");
4086 assert_eq!(frame.size(), 0);
4087 assert_eq!(frame.local(0), Some(-8));
4088 }
4089
4090 /// An `alloca` whose size is an operand, which is a variable length array.
4091 fn growing(source: &mut Func, block: Block, size: Value, align: u32) -> Value {
4092 let info = MemInfo { size: 0, align, ..plain() };
4093 let mut build = Builder::new(source, block);
4094 let mem = build.func().add_mem(info);
4095 let args = build.func().push_values(&[size]);
4096 build.value(
4097 InstData { args, extra: Extra::Mem(mem), ..InstData::new(Opcode::Alloca) },
4098 Type::PTR,
4099 )
4100 }
4101
4102 #[test]
4103 fn a_stack_slot_whose_size_is_not_known_until_it_runs_takes_the_bytes_off_the_stack_pointer() {
4104 let (mut names, mut source, block, args) = blank(&[Type::int(64)]);
4105 let slot = growing(&mut source, block, args[0], 16);
4106 Builder::new(&mut source, block).ret(&[slot]);
4107
4108 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4109 .expect("every instruction has a rule");
4110
4111 // The bytes come off the stack pointer where the declaration stands and the address is
4112 // where the stack pointer then is, which is one subtraction and one `lea` rather than a
4113 // slot the frame laid out. Nothing is on the list of locals, because there is nothing
4114 // about this the frame could place.
4115 let text = mir::print_func(&lowered.func, &names, ®S);
4116 assert!(text.contains("$rsp = x64.sub_rr_64 $rsp, %0"), "{text}");
4117 assert!(text.contains("x64.lea_64 [$rsp]"), "{text}");
4118 assert!(lowered.stack.locals.is_empty(), "{text}");
4119 assert_eq!(lowered.stack.dynamic.len(), 1);
4120 assert!(lowered.stack.grown_at.is_some());
4121 }
4122
4123 #[test]
4124 fn a_growing_slot_wanting_more_alignment_than_the_stack_pointer_has_is_reported() {
4125 let (mut names, mut source, block, args) = blank(&[Type::int(64)]);
4126 let slot = growing(&mut source, block, args[0], 32);
4127 Builder::new(&mut source, block).ret(&[slot]);
4128
4129 // Thirty two is more than a call leaves the stack pointer on, so giving it what it asked
4130 // for means masking the stack pointer after moving it, and after that no constant reaches
4131 // the rest of the frame from the frame pointer either. A second pointer held for the
4132 // purpose is what fixes it and there is not one yet.
4133 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4134 .expect_err("nothing realigns a frame that grows");
4135 assert_eq!(
4136 failed.to_string(),
4137 "this local wants more alignment than the stack pointer is left on, which needs a \
4138 base register nothing here keeps"
4139 );
4140 }
4141
4142 #[test]
4143 fn a_frame_that_grows_reaches_its_own_locals_through_the_frame_pointer() {
4144 let (mut names, mut source, block, args) = blank(&[Type::int(64)]);
4145 let fixed = slot(&mut source, block, 4, 4);
4146 let mut build = Builder::new(&mut source, block);
4147 let nine = build.iconst(Type::int(32), 9);
4148 build.store(nine, fixed, plain(), Flags::default());
4149 let grown = growing(&mut source, block, args[0], 16);
4150 Builder::new(&mut source, block).ret(&[grown]);
4151
4152 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4153 .expect("every instruction has a rule");
4154 let stack = lowered.stack;
4155 let mut out = lowered.func;
4156 let env = env();
4157 let allocation = rucc_regalloc::run(&mut out, &env, "test");
4158 let layout = stack.layout(Layout::new(&SYSV, REGS));
4159 let frame = Frame::of(&out, &allocation, &layout);
4160 finish(&mut out, &allocation, &frame, &stack, Convention::new(&SYSV, &FRAME), &mut names);
4161
4162 // The stack pointer moves in the middle of the function, so the four bytes of the fixed
4163 // local are not a constant away from it any more and the frame pointer is what reaches
4164 // them. The frame keeps one whatever the flags asked for, takes its bytes rather than
4165 // living in the red zone, and the address of the growing slot is off the stack pointer as
4166 // it stands after the subtraction rather than off anything the prologue left.
4167 let text = mir::print_func(&out, &names, ®S);
4168 assert!(frame.grows());
4169 assert!(frame.frame_pointer());
4170 assert!(frame.size() > 0, "{text}");
4171 assert!(text.contains("x64.lea_64 [$rbp"), "{text}");
4172 assert!(text.contains("$rsp = x64.sub_rr_64 $rsp"), "{text}");
4173 assert!(text.contains("x64.lea_64 [$rsp]"), "{text}");
4174 }
4175
4176 #[test]
4177 fn an_address_is_read_written_and_added_to_like_the_integer_it_is() {
4178 let (mut names, mut source, block, args) = blank(&[Type::PTR, Type::int(64)]);
4179 let mut build = Builder::new(&mut source, block);
4180 let stepped = build.func().push_values(&[args[0], args[1]]);
4181 let next =
4182 build.value(InstData { args: stepped, ..InstData::new(Opcode::PtrAdd) }, Type::PTR);
4183 let loaded = build.load(Type::int(32), next, plain(), Flags::default());
4184 build.ret(&[loaded]);
4185
4186 // `int f(int *p, long i) { return *(int *)((char *)p + i); }`. Nothing about this is new
4187 // in the rule set, which is the point: the two addresses arrive in registers because an
4188 // address is an integer as wide as one, and the arithmetic on them is the add it always
4189 // was, so every rule written about an add reaches it.
4190 //
4191 // The add stays its own instruction rather than folding into the address the load reads
4192 // from. Two registers with no scale on either is the one addressing mode the rules have no
4193 // load through, because the folds that exist are the displacement one and the scaled ones,
4194 // and this is neither. That is a peephole worth having and not a thing this changes.
4195 assert_eq!(
4196 lower(&mut names, &source),
4197 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
4198 %1:gpr($rsi) = x64.arg_val_64\n %2:gpr(reuse 1) = x64.add_rr_64 %0, %1\n \
4199 %3:gpr = x64.mov_rm_32 [%2]\n x64.ret_val_32 %3($rax)\n}\n"
4200 );
4201 }
4202
4203 /// The address of a file scope name, which is what every use of a global and every string
4204 /// literal starts from.
4205 fn address_of(source: &mut Func, block: Block, names: &mut Interner, name: &str) -> Value {
4206 let symbol = names.intern(name);
4207 let mut build = Builder::new(source, block);
4208 build.value(
4209 InstData { extra: Extra::Symbol(symbol), ..InstData::new(Opcode::GlobalAddr) },
4210 Type::PTR,
4211 )
4212 }
4213
4214 #[test]
4215 fn the_address_of_a_name_is_one_instruction_carrying_the_name() {
4216 let (mut names, mut source, block, _) = blank(&[]);
4217 let counter = address_of(&mut source, block, &mut names, "counter");
4218 let mut build = Builder::new(&mut source, block);
4219 let loaded = build.load(Type::int(32), counter, plain(), Flags::default());
4220 build.ret(&[loaded]);
4221
4222 // `extern int counter; int f(void) { return counter; }`. The address is an addressing mode
4223 // that names no register and carries the symbol, which is what the assembler writes
4224 // relative to `%rip` and what the object writer leaves a relocation for.
4225 assert_eq!(
4226 lower(&mut names, &source),
4227 "mfunc @f {\nblock0:\n %0:gpr = x64.lea_64 [@counter]\n \
4228 %1:gpr = x64.mov_rm_32 [%0]\n x64.ret_val_32 %1($rax)\n}\n"
4229 );
4230 }
4231
4232 #[test]
4233 fn the_address_of_a_name_outside_the_file_is_read_out_of_the_offset_table() {
4234 let (mut names, mut source, block, _) = blank(&[]);
4235 let away = address_of(&mut source, block, &mut names, "away");
4236 Builder::new(&mut source, block).ret(&[away]);
4237 let elsewhere: Elsewhere = [names.intern("away")].into_iter().collect();
4238
4239 // `extern void away(void); void *f(void) { return away; }`. A load and not an address
4240 // computation, because the distance from here to a name a shared library may be the one
4241 // that defines is not a number any link can work out, and the slot the linker fills in is
4242 // in this program and so is a distance it has.
4243 let out =
4244 func(&source, &mut names, &SYSV, &elsewhere).expect("every instruction has a rule");
4245 assert_eq!(
4246 mir::print_func(&out.func, &names, ®S),
4247 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_rm_64 [got @away]\n \
4248 x64.ret_val_64 %0($rax)\n}\n"
4249 );
4250 }
4251
4252 #[test]
4253 fn the_address_of_a_thread_local_is_an_offset_out_of_the_table_plus_where_this_thread_starts() {
4254 let (mut names, mut source, block, _) = blank(&[]);
4255 let own = address_of(&mut source, block, &mut names, "own");
4256 Builder::new(&mut source, block).ret(&[own]);
4257 let elsewhere = Elsewhere::default().with_threads([names.intern("own")]);
4258
4259 // `extern _Thread_local int own; void *f(void) { return &own; }`. Three instructions where
4260 // the two cases above are one, because there is no address to load or to work out: the
4261 // slot holds how far into a thread's block the variable sits, `%fs:0` is where this
4262 // thread's block starts, and the sum of the two is this thread's copy.
4263 let out =
4264 func(&source, &mut names, &SYSV, &elsewhere).expect("every instruction has a rule");
4265 assert_eq!(
4266 mir::print_func(&out.func, &names, ®S),
4267 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_rm_64 [thread @own]\n \
4268 %1:gpr = x64.mov_rm_64 [fs:0]\n %2:gpr(reuse 1) = x64.add_rr_64 %0, %1\n \
4269 x64.ret_val_64 %2($rax)\n}\n"
4270 );
4271 }
4272
4273 /// The same load with nothing added to it, which is the whole of `__builtin_thread_pointer`.
4274 #[test]
4275 fn the_start_of_this_thread_s_own_storage_is_the_one_load_and_no_arithmetic() {
4276 let (mut names, mut source, block, _) = blank(&[]);
4277 let here =
4278 Builder::new(&mut source, block).value(InstData::new(Opcode::ThreadPointer), Type::PTR);
4279 Builder::new(&mut source, block).ret(&[here]);
4280
4281 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
4282 .expect("every instruction has a rule");
4283 assert_eq!(
4284 mir::print_func(&out.func, &names, ®S),
4285 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_rm_64 [fs:0]\n \
4286 x64.ret_val_64 %0($rax)\n}\n"
4287 );
4288 }
4289
4290 /// One `asm` statement, with its template and its constraint list written as a program does.
4291 fn assembly(
4292 source: &mut Func,
4293 block: Block,
4294 names: &mut Interner,
4295 template: &str,
4296 constraints: &str,
4297 args: &[Value],
4298 results: &[Type],
4299 ) -> Inst {
4300 let info = AsmInfo {
4301 template: names.intern(template),
4302 constraints: names.intern(constraints),
4303 clobbers: names.intern("memory"),
4304 targets: rucc_ir::BlockCallList::EMPTY,
4305 };
4306 Builder::new(source, block).inline_asm(info, args, results, Flags::VOLATILE)
4307 }
4308
4309 #[test]
4310 fn an_asm_with_an_empty_template_and_no_operands_is_no_instructions() {
4311 let (mut names, mut source, block, _) = blank(&[]);
4312 assembly(&mut source, block, &mut names, "", "", &[], &[]);
4313 Builder::new(&mut source, block).ret(&[]);
4314
4315 // `asm volatile ("" : : : "memory")`, which is a barrier and nothing else. The barrier was
4316 // spent on the optimizer, which has finished by now, so what is left is nothing.
4317 assert_eq!(lower(&mut names, &source), "mfunc @f {\nblock0:\n}\n");
4318 }
4319
4320 #[test]
4321 fn an_output_an_input_is_tied_to_is_the_register_that_input_arrived_in() {
4322 let i32 = Type::int(32);
4323 let (mut names, mut source, block, args) = blank(&[i32]);
4324 let out = assembly(&mut source, block, &mut names, "", "=r,0", &args, &[i32]);
4325 let produced = source[out].results().next().expect("one result");
4326 Builder::new(&mut source, block).ret(&[produced]);
4327
4328 // `asm ("" : "=r" (x) : "0" (x))`, which is how a program stops the optimizer following a
4329 // value without changing it. The two share a place and the template writes nothing over
4330 // it, so the value comes back out of the register it went in.
4331 assert_eq!(
4332 lower(&mut names, &source),
4333 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
4334 x64.ret_val_32 %0($rax)\n}\n"
4335 );
4336 }
4337
4338 #[test]
4339 fn an_output_written_plus_is_the_same_rename() {
4340 let i32 = Type::int(32);
4341 let (mut names, mut source, block, args) = blank(&[i32]);
4342 let out = assembly(&mut source, block, &mut names, "", "+r", &args, &[i32]);
4343 let produced = source[out].results().next().expect("one result");
4344 Builder::new(&mut source, block).ret(&[produced]);
4345
4346 // `asm ("" : "+r" (x))`, which says the same thing in one operand instead of two.
4347 assert_eq!(
4348 lower(&mut names, &source),
4349 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
4350 x64.ret_val_32 %0($rax)\n}\n"
4351 );
4352 }
4353
4354 #[test]
4355 fn an_output_nothing_is_tied_to_is_a_zero() {
4356 let i32 = Type::int(32);
4357 let (mut names, mut source, block, _) = blank(&[]);
4358 let out = assembly(&mut source, block, &mut names, "", "=r", &[], &[i32]);
4359 let produced = source[out].results().next().expect("one result");
4360 Builder::new(&mut source, block).ret(&[produced]);
4361
4362 // `asm ("" : "=r" (y))`, whose answer is whatever the assembly left in the register, and
4363 // an empty template leaves nothing. A definite value rather than a register nothing wrote,
4364 // because the allocator is owed a definition before the use however little the program is.
4365 assert_eq!(
4366 lower(&mut names, &source),
4367 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_ri_32 0\n x64.ret_val_32 %0($rax)\n}\n"
4368 );
4369 }
4370
4371 #[test]
4372 fn an_asm_with_instructions_in_its_template_is_refused_as_an_asm() {
4373 let (mut names, mut source, block, _) = blank(&[]);
4374 assembly(&mut source, block, &mut names, "nop", "", &[], &[]);
4375 Builder::new(&mut source, block).ret(&[]);
4376
4377 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4378 .expect_err("nothing here assembles a template");
4379 assert_eq!(
4380 failed.to_string(),
4381 "this `asm` has instructions in its template, which nothing here assembles"
4382 );
4383 }
4384
4385 #[test]
4386 fn a_constraint_list_that_does_not_describe_the_operands_is_refused() {
4387 let i32 = Type::int(32);
4388 let (mut names, mut source, block, args) = blank(&[i32]);
4389 assembly(&mut source, block, &mut names, "", "=r", &args, &[]);
4390 Builder::new(&mut source, block).ret(&[]);
4391
4392 // An output with no result to be, which is what the front end never writes and what a
4393 // hand written module can. Refused rather than placed by a guess.
4394 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4395 .expect_err("the list and the instruction disagree");
4396 assert_eq!(failed.to_string(), "this `asm` has an operand this cannot place");
4397 }
4398
4399 /// A cast between a pointer and an integer, at whatever width the result is asked for.
4400 fn cast(source: &mut Func, block: Block, opcode: Opcode, from: Value, to: Type) -> Value {
4401 let mut build = Builder::new(source, block);
4402 let args = build.func().push_values(&[from]);
4403 build.value(InstData { args, ..InstData::new(opcode) }, to)
4404 }
4405
4406 #[test]
4407 fn a_cast_between_a_pointer_and_an_integer_as_wide_is_no_instruction_at_all() {
4408 let (mut names, mut source, block, args) = blank(&[Type::PTR]);
4409 let number = cast(&mut source, block, Opcode::PtrToInt, args[0], Type::int(64));
4410 Builder::new(&mut source, block).ret(&[number]);
4411
4412 // `long f(void *p) { return (long)p; }`. An address on this machine is an integer as wide
4413 // as the machine addresses, so the cast changes what the type system calls the value and
4414 // changes nothing about the value, and the register holding it is the one that held it.
4415 assert_eq!(
4416 lower(&mut names, &source),
4417 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
4418 x64.ret_val_64 %0($rax)\n}\n"
4419 );
4420 }
4421
4422 #[test]
4423 fn a_null_pointer_is_a_constant_that_reaches_a_register_before_anything_reads_it() {
4424 let (mut names, mut source, block, _) = blank(&[]);
4425 let mut build = Builder::new(&mut source, block);
4426 let zero = build.iconst(Type::int(64), 0);
4427 let null = cast(&mut source, block, Opcode::IntToPtr, zero, Type::PTR);
4428 Builder::new(&mut source, block).ret(&[null]);
4429
4430 // `void *f(void) { return 0; }`. The cast is nothing, and reading its operand is what
4431 // writes the zero down: a constant is materialized where it is wanted rather than where
4432 // the IR defined it, and without the read there would be no instruction at all.
4433 assert_eq!(
4434 lower(&mut names, &source),
4435 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_ri_64 0\n x64.ret_val_64 %0($rax)\n}\n"
4436 );
4437 }
4438
4439 #[test]
4440 fn the_five_linkages_the_ir_has_narrow_to_the_three_an_object_file_can_say() {
4441 let readings = [
4442 (Linkage::External, mir::Binding::Global),
4443 (Linkage::Common, mir::Binding::Global),
4444 (Linkage::Internal, mir::Binding::Local),
4445 (Linkage::Weak, mir::Binding::Weak),
4446 (Linkage::LinkOnce, mir::Binding::Weak),
4447 ];
4448 for (linkage, wanted) in readings {
4449 let (mut names, mut source, block, _) = blank(&[]);
4450 source.linkage = linkage;
4451 Builder::new(&mut source, block).ret(&[]);
4452 let out = func(&source, &mut names, &SYSV, &Elsewhere::default()).expect("a return");
4453 // The narrowing is done here rather than where the object is written, because a
4454 // machine function is all the assembler and the writer are ever handed.
4455 assert_eq!(out.func.binding, wanted, "{linkage:?}");
4456 }
4457 }
4458
4459 /// The visibility makes the same trip and is not narrowed on the way, because ELF says all
4460 /// three of them.
4461 ///
4462 /// Here for the reason the linkage above is here. A machine function is the whole of what the
4463 /// assembler and the object writer are handed, so a fact about the symbol that does not get
4464 /// onto one is a fact that is gone by the time anything could write it down, and the way that
4465 /// shows up is a shared library exporting the wrong set of names with nothing said anywhere.
4466 #[test]
4467 fn the_visibility_survives_the_trip_from_the_ir_to_a_machine_function() {
4468 let readings = [
4469 (Visibility::Default, mir::Visibility::Default),
4470 (Visibility::Hidden, mir::Visibility::Hidden),
4471 (Visibility::Protected, mir::Visibility::Protected),
4472 ];
4473 for (visibility, wanted) in readings {
4474 let (mut names, mut source, block, _) = blank(&[]);
4475 source.visibility = visibility;
4476 Builder::new(&mut source, block).ret(&[]);
4477 let out = func(&source, &mut names, &SYSV, &Elsewhere::default()).expect("a return");
4478 assert_eq!(out.func.visibility, wanted, "{visibility:?}");
4479 }
4480 }
4481
4482 #[test]
4483 fn a_cast_between_a_pointer_and_a_narrower_integer_is_reported() {
4484 let (mut names, mut source, block, args) = blank(&[Type::PTR]);
4485 let number = cast(&mut source, block, Opcode::PtrToInt, args[0], Type::int(32));
4486 Builder::new(&mut source, block).ret(&[number]);
4487
4488 // The front end never writes one: it casts at the address width and truncates or extends
4489 // around it, so both of those are the rules they always were. IR from somewhere else that
4490 // does write one is refused rather than compiled to a move that keeps the high half.
4491 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4492 .expect_err("no rule narrows an address");
4493 assert_eq!(failed.to_string(), "no rule lowers a `ptrtoint` producing a `i32`");
4494 }
4495
4496 /// The type this machine has no register for.
4497 fn long_double() -> Type {
4498 Type::float(rucc_ir::Float::F80)
4499 }
4500
4501 #[test]
4502 fn a_double_widened_and_narrowed_again_goes_out_through_the_frame_and_back() {
4503 let f64 = Type::float(rucc_ir::Float::F64);
4504 let (mut names, mut source, block, args) = blank(&[f64]);
4505 let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
4506 let back = cast(&mut source, block, Opcode::FPTrunc, wide, f64);
4507 Builder::new(&mut source, block).ret(&[back]);
4508
4509 // `double f(double d) { long double x = d; return x; }`. The x87 reads memory and nothing
4510 // else, so the value is written to the crossing slot, loaded at the format that widens it
4511 // and put in the slot the eighty bit value lives in. Coming back is the same three the
4512 // other way. Both slots are addressed by a `lea` with nothing in it yet, which is what
4513 // every address in a frame looks like here until `finish` has the numbers.
4514 assert_eq!(
4515 lower(&mut names, &source),
4516 "mfunc @f {\nblock0:\n \
4517 %0:xmm($xmm0) = x64.arg_val_f64\n \
4518 %1:gpr = x64.lea_64 [$rsp]\n \
4519 %2:gpr = x64.lea_64 [$rsp]\n \
4520 x64.movsd_mr %0, [%1]\n \
4521 x64.fld_l [%1]\n \
4522 x64.fstp_t [%2]\n \
4523 %3:gpr = x64.lea_64 [$rsp]\n \
4524 %4:gpr = x64.lea_64 [$rsp]\n \
4525 x64.fld_t [%3]\n \
4526 x64.fstp_l [%4]\n \
4527 %5:xmm = x64.movsd_rm [%4]\n \
4528 x64.ret_val_f64 %5($xmm0)\n}\n"
4529 );
4530 }
4531
4532 #[test]
4533 fn a_long_double_has_sixteen_bytes_of_its_own_and_keeps_them() {
4534 let f64 = Type::float(rucc_ir::Float::F64);
4535 let (mut names, mut source, block, args) = blank(&[f64]);
4536 let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
4537 let once = cast(&mut source, block, Opcode::FPTrunc, wide, f64);
4538 let twice = cast(&mut source, block, Opcode::FPTrunc, wide, f64);
4539 let mut build = Builder::new(&mut source, block);
4540 let sum = build.binary(Opcode::FAdd, once, twice, Flags::default());
4541 build.ret(&[sum]);
4542
4543 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
4544 .expect("every instruction is written");
4545
4546 // Two slots and not four: sixteen bytes for the one eighty bit value, which is what the
4547 // psABI says one takes and is aligned to, and eight for the crossing, which every group
4548 // in the function shares because nothing is ever left in it. The value's slot is its own
4549 // for the whole function, so reading it twice reads the same sixteen bytes.
4550 assert_eq!(
4551 out.stack.locals,
4552 vec![Local { size: 8, align: 8 }, Local { size: 16, align: 16 }]
4553 );
4554 }
4555
4556 #[test]
4557 fn an_integer_becomes_a_long_double_by_being_loaded_as_one() {
4558 let (mut names, mut source, block, args) = blank(&[Type::int(64)]);
4559 let wide = cast(&mut source, block, Opcode::SIToFP, args[0], long_double());
4560 let back =
4561 cast(&mut source, block, Opcode::FPTrunc, wide, Type::float(rucc_ir::Float::F64));
4562 Builder::new(&mut source, block).ret(&[back]);
4563
4564 // `double f(long n) { long double x = n; return x; }`. `fild` is the same push at another
4565 // format, so the conversion is the load and there is no instruction that converts.
4566 let text = lower(&mut names, &source);
4567 assert!(text.contains("x64.mov_mr_64 %0, [%1]"), "{text}");
4568 assert!(text.contains("x64.fild_ll [%1]"), "{text}");
4569 }
4570
4571 #[test]
4572 fn a_long_double_becoming_an_integer_cuts_towards_zero_with_the_control_word() {
4573 let (mut names, mut source, block, args) = blank(&[Type::float(rucc_ir::Float::F64)]);
4574 let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
4575 let whole = cast(&mut source, block, Opcode::FPToSI, wide, Type::int(32));
4576 Builder::new(&mut source, block).ret(&[whole]);
4577
4578 // The one conversion here with no single instruction behind it. C cuts towards zero and
4579 // the unit rounds the way its control word says, so the word is saved, ORed with the two
4580 // bits that mean truncate, loaded, used and put back. Nine instructions for what `fisttp`
4581 // does in one, and `spec/10-backend.md` section 10.8 says why that one is not used.
4582 let text = lower(&mut names, &source);
4583 let group: Vec<&str> = text
4584 .lines()
4585 .map(str::trim)
4586 .filter(|line| line.starts_with("x64.f") || line.contains("_16"))
4587 .collect();
4588 assert_eq!(
4589 group,
4590 [
4591 "x64.fld_l [%1]",
4592 "x64.fstp_t [%2]",
4593 "x64.fnstcw [%5]",
4594 "%6:gpr = x64.mov_rm_16 [%5]",
4595 "%7:gpr(reuse 1) = x64.or_ri_16 %6, 3072",
4596 "x64.mov_mr_16 %7, [%5 + 2]",
4597 "x64.fldcw [%5 + 2]",
4598 "x64.fld_t [%3]",
4599 "x64.fistp_l [%4]",
4600 "x64.fldcw [%5]",
4601 ],
4602 "{text}"
4603 );
4604 }
4605
4606 #[test]
4607 fn a_long_double_is_read_and_written_as_the_bits_it_already_is() {
4608 let (mut names, mut source, block, args) = blank(&[Type::PTR, Type::PTR]);
4609 let mut build = Builder::new(&mut source, block);
4610 let value = build.load(long_double(), args[0], plain(), Flags::default());
4611 build.store(value, args[1], plain(), Flags::default());
4612 build.ret(&[]);
4613
4614 // `void f(long double *a, long double *b) { *b = *a; }`. A copy is a push and a pop at the
4615 // format the value is already in, which neither converts nor looks: a signalling NaN stays
4616 // one and nothing is raised, which is the whole of what makes it a copy.
4617 let text = lower(&mut names, &source);
4618 let group: Vec<&str> =
4619 text.lines().map(str::trim).filter(|line| line.starts_with("x64.f")).collect();
4620 assert_eq!(
4621 group,
4622 ["x64.fld_t [%0]", "x64.fstp_t [%2]", "x64.fld_t [%3]", "x64.fstp_t [%1]"],
4623 "{text}"
4624 );
4625 }
4626
4627 /// Two `long double` values, from two `double` parameters, and the instructions that made
4628 /// them, which every test below this one throws away.
4629 fn two_long_doubles(source: &mut Func, block: Block, args: &[Value]) -> (Value, Value) {
4630 let left = cast(source, block, Opcode::FPExt, args[0], long_double());
4631 let right = cast(source, block, Opcode::FPExt, args[1], long_double());
4632 (left, right)
4633 }
4634
4635 /// The x87 instructions of a function, in order, with everything else dropped.
4636 fn stack_only(text: &str) -> Vec<&str> {
4637 text.lines().map(str::trim).filter(|line| line.contains("x64.f")).collect()
4638 }
4639
4640 /// The two frame slots the last two addresses of a function were taken of, which in a
4641 /// comparison are the two operands in the order they go on the stack.
4642 fn pushed(out: &Lowered) -> Vec<usize> {
4643 let taken: Vec<usize> = out.stack.addresses.iter().map(|&(_, local)| local).collect();
4644 taken[taken.len() - 2..].to_vec()
4645 }
4646
4647 #[test]
4648 fn adding_two_long_doubles_pushes_both_and_leaves_the_answer_in_a_slot() {
4649 let f64 = Type::float(rucc_ir::Float::F64);
4650 let (mut names, mut source, block, args) = blank(&[f64, f64]);
4651 let (left, right) = two_long_doubles(&mut source, block, &args);
4652 let sum =
4653 Builder::new(&mut source, block).binary(Opcode::FAdd, left, right, Flags::default());
4654 let back = cast(&mut source, block, Opcode::FPTrunc, sum, f64);
4655 Builder::new(&mut source, block).ret(&[back]);
4656
4657 // `double f(double a, double b) { return (long double) a + (long double) b; }`. The last
4658 // four lines are the add: both operands pushed, the instruction that names neither of
4659 // them because they are the top two of a stack, and the answer taken off into its slot.
4660 let text = lower(&mut names, &source);
4661 assert_eq!(
4662 stack_only(&text),
4663 [
4664 "x64.fld_l [%2]",
4665 "x64.fstp_t [%3]",
4666 "x64.fld_l [%4]",
4667 "x64.fstp_t [%5]",
4668 "x64.fld_t [%6]",
4669 "x64.fld_t [%7]",
4670 "x64.fadd_p",
4671 "x64.fstp_t [%8]",
4672 "x64.fld_t [%9]",
4673 "x64.fstp_l [%10]",
4674 ],
4675 "{text}"
4676 );
4677 }
4678
4679 #[test]
4680 fn a_subtraction_pushes_the_left_operand_first_and_asks_for_the_att_spelling() {
4681 let f64 = Type::float(rucc_ir::Float::F64);
4682 let (mut names, mut source, block, args) = blank(&[f64, f64]);
4683 let (left, right) = two_long_doubles(&mut source, block, &args);
4684 let less =
4685 Builder::new(&mut source, block).binary(Opcode::FSub, left, right, Flags::default());
4686 let back = cast(&mut source, block, Opcode::FPTrunc, less, f64);
4687 Builder::new(&mut source, block).ret(&[back]);
4688
4689 // The left one goes on first, so it ends up under the right one, and the answer wanted is
4690 // the one below minus the top. In AT&T that is `fsubrp`, since `fsubp` there is `DE E0+i`
4691 // and computes the other one. The `r` says which spelling this is and not which order the
4692 // pushes were in. `crates/rucc/tests/x87.rs` is what says the answer is right, because a
4693 // name is what got this wrong the first time.
4694 let text = lower(&mut names, &source);
4695 assert_eq!(
4696 &stack_only(&text)[4..8],
4697 ["x64.fld_t [%6]", "x64.fld_t [%7]", "x64.fsubr_p", "x64.fstp_t [%8]"],
4698 "{text}"
4699 );
4700 }
4701
4702 #[test]
4703 fn negating_a_long_double_turns_the_sign_over_and_reads_nothing() {
4704 let f64 = Type::float(rucc_ir::Float::F64);
4705 let (mut names, mut source, block, args) = blank(&[f64]);
4706 let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
4707 let flipped = Builder::new(&mut source, block).unary(Opcode::FNeg, wide, long_double());
4708 let back = cast(&mut source, block, Opcode::FPTrunc, flipped, f64);
4709 Builder::new(&mut source, block).ret(&[back]);
4710
4711 // `fchs` and not a subtraction from zero, which would give a different answer at a negative
4712 // zero and would signal at a NaN. It does not read the value as a number at all.
4713 let text = lower(&mut names, &source);
4714 assert_eq!(
4715 &stack_only(&text)[2..5],
4716 ["x64.fld_t [%3]", "x64.fchs", "x64.fstp_t [%4]"],
4717 "{text}"
4718 );
4719 }
4720
4721 #[test]
4722 fn comparing_two_long_doubles_puts_the_left_one_on_top() {
4723 let f64 = Type::float(rucc_ir::Float::F64);
4724 let (mut names, mut source, block, args) = blank(&[f64, f64]);
4725 let (left, right) = two_long_doubles(&mut source, block, &args);
4726 let mut build = Builder::new(&mut source, block);
4727 build.fcmp(FloatPred::Ogt, left, right, Flags::default());
4728 build.ret(&[]);
4729
4730 // `a > b`. `fucomip` asks about the top of the stack against what is under it, so the
4731 // operand the predicate is about has to go on last, which is the other way round from the
4732 // arithmetic above. The pop that clears the loser and the byte that reads the flags are
4733 // both inside the one opcode.
4734 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
4735 .expect("every instruction is written");
4736 let slots = pushed(&out);
4737 assert_eq!(slots, [2, 1], "the right operand goes on first and the left one on top");
4738 let text = mir::print_func(&out.func, &names, ®S);
4739 assert_eq!(
4740 &stack_only(&text)[4..],
4741 ["x64.fld_t [%6]", "x64.fld_t [%7]", "%8:gpr = x64.fucomip_set_a"],
4742 "{text}"
4743 );
4744 }
4745
4746 #[test]
4747 fn a_comparison_that_the_machine_has_backwards_swaps_the_two_pushes() {
4748 let f64 = Type::float(rucc_ir::Float::F64);
4749 let (mut names, mut source, block, args) = blank(&[f64, f64]);
4750 let (left, right) = two_long_doubles(&mut source, block, &args);
4751 let mut build = Builder::new(&mut source, block);
4752 build.fcmp(FloatPred::Olt, left, right, Flags::default());
4753 build.ret(&[]);
4754
4755 // `a < b` is `b > a` and this machine has the one condition, so the same opcode runs with
4756 // the operands the other way round. The same trade the vector rules make, and it has to
4757 // be the same one: a `long double` comparison that picked a different condition from the
4758 // `double` comparison of the same two numbers would be wrong at exactly the unordered
4759 // cases the two conditions differ on.
4760 //
4761 // Which slot each push names is the whole of the difference from the test above, and the
4762 // text does not show it, since an address in a frame is a `lea` with nothing in it until
4763 // `finish` has the numbers. So the slots are what is read here.
4764 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
4765 .expect("every instruction is written");
4766 let slots = pushed(&out);
4767 assert_eq!(slots, [1, 2], "the left operand goes on first and the right one on top");
4768 let text = mir::print_func(&out.func, &names, ®S);
4769 assert_eq!(
4770 &stack_only(&text)[4..],
4771 ["x64.fld_t [%6]", "x64.fld_t [%7]", "%8:gpr = x64.fucomip_set_a"],
4772 "{text}"
4773 );
4774 }
4775
4776 #[test]
4777 fn an_ordered_equal_needs_a_second_byte_to_put_the_two_conditions_together() {
4778 let f64 = Type::float(rucc_ir::Float::F64);
4779 let (mut names, mut source, block, args) = blank(&[f64, f64]);
4780 let (left, right) = two_long_doubles(&mut source, block, &args);
4781 let mut build = Builder::new(&mut source, block);
4782 build.fcmp(FloatPred::Oeq, left, right, Flags::default());
4783 build.ret(&[]);
4784
4785 // Equal and ordered are two conditions and the flags carry both, so the opcode writes a
4786 // second register as well as the one the value is in and ANDs them together. Said here by
4787 // handing it a spare, since an instruction that wrote a register nothing knew about would
4788 // be an instruction the allocator could put a live value in the way of.
4789 let text = lower(&mut names, &source);
4790 assert!(text.contains("%8:gpr, %9:gpr = x64.fucomip_set_e_and_np"), "{text}");
4791 }
4792
4793 #[test]
4794 fn a_comparison_that_is_never_asked_is_reported() {
4795 let f64 = Type::float(rucc_ir::Float::F64);
4796 let (mut names, mut source, block, args) = blank(&[f64, f64]);
4797 let (left, right) = two_long_doubles(&mut source, block, &args);
4798 let mut build = Builder::new(&mut source, block);
4799 build.fcmp(FloatPred::False, left, right, Flags::default());
4800 build.ret(&[]);
4801
4802 // Always false is a constant and not a comparison, so there is no condition to pick and
4803 // nothing here folds it into one: an instruction that quietly agreed with it would hide
4804 // that the optimizer left a comparison in that it should have taken out.
4805 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4806 .expect_err("no condition is always false");
4807 assert_eq!(failed.to_string(), "no rule lowers a `fcmp` producing a `i1`");
4808 }
4809
4810 #[test]
4811 fn a_long_double_constant_is_the_bits_of_it_put_where_the_value_lives() {
4812 let (mut names, mut source, block, args) = blank(&[Type::PTR]);
4813 let mut build = Builder::new(&mut source, block);
4814 // `1.5L`, which is the leading bit and one more of significand, and an exponent of zero.
4815 let one_and_a_half = build.fconst(long_double(), 0x3fff_c000_0000_0000_0000);
4816 build.store(one_and_a_half, args[0], plain(), Flags::default());
4817 build.ret(&[]);
4818
4819 // No x87 instruction at all. A slot holding one of these is the value, so a constant is
4820 // its ten bytes written where the value lives, and whatever reads it does the `fld`.
4821 let text = lower(&mut names, &source);
4822 assert!(text.contains("x64.mov_ri_64 -4611686018427387904"), "{text}");
4823 assert!(text.contains("x64.mov_ri_16 16383"), "{text}");
4824 assert!(text.contains("x64.mov_mr_16 %3, [%1 + 8]"), "{text}");
4825 // The six bytes above the ten are the padding that makes the type sixteen wide, and they
4826 // are unspecified rather than zero, so nothing writes them.
4827 assert_eq!(text.matches("x64.mov_mr").count(), 2, "{text}");
4828 }
4829
4830 #[test]
4831 fn a_negative_long_double_constant_keeps_the_bit_above_its_exponent() {
4832 let (mut names, mut source, block, args) = blank(&[Type::PTR]);
4833 let mut build = Builder::new(&mut source, block);
4834 let minus = build.fconst(long_double(), 0xbfff_c000_0000_0000_0000);
4835 build.store(minus, args[0], plain(), Flags::default());
4836 build.ret(&[]);
4837
4838 // `-1.5L`. The sign is the top bit of the two byte half, so the immediate that half is put
4839 // in a register with is above the signed range of sixteen bits and has to stay there: read
4840 // as a number it would be negative, and it is not a number, it is two bytes.
4841 let text = lower(&mut names, &source);
4842 assert!(text.contains("x64.mov_ri_16 49151"), "{text}");
4843 }
4844
4845 #[test]
4846 fn a_long_double_crosses_an_edge_as_an_address_and_is_copied_where_it_lands() {
4847 let (mut names, mut source, block, args) = blank(&[Type::float(rucc_ir::Float::F64)]);
4848 let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
4849 let next = source.create_block();
4850 let param = source.append_param(next, long_double());
4851 Builder::new(&mut source, block).jump(next, &[wide]);
4852 Builder::new(&mut source, next).ret(&[param]);
4853
4854 // What the edge carries is the address of the slot the value is already in, which is an
4855 // ordinary register the allocator has an opinion about. The block on the other side copies
4856 // the sixteen bytes into a slot of its own before anything reads them, so a second edge
4857 // handing over a second address would still leave one place for a reader to look.
4858 let text = lower(&mut names, &source);
4859 let second: Vec<&str> = text
4860 .lines()
4861 .skip_while(|line| !line.starts_with("block1"))
4862 .skip(1)
4863 .take(3)
4864 .map(str::trim)
4865 .collect();
4866 assert_eq!(
4867 second,
4868 ["x64.fld_t [%4]", "%5:gpr = x64.lea_64 [$rsp]", "x64.fstp_t [%5]"],
4869 "{text}"
4870 );
4871 }
4872
4873 #[test]
4874 fn more_long_doubles_at_a_block_than_the_stack_is_deep_are_reported() {
4875 let f64 = Type::float(rucc_ir::Float::F64);
4876 let (mut names, mut source, block, args) = blank(&[f64]);
4877 let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
4878 let next = source.create_block();
4879 let params: Vec<Value> =
4880 (0..=X87_DEPTH).map(|_| source.append_param(next, long_double())).collect();
4881 let carried: Vec<Value> = params.iter().map(|_| wide).collect();
4882 Builder::new(&mut source, block).jump(next, &carried);
4883 Builder::new(&mut source, next).ret(&[params[0]]);
4884
4885 // The copies go through the x87 stack so that every one of them is read before any of them
4886 // is written, which is what makes a block that swaps two of these right. Nine of them do
4887 // not fit on the stack, and copying the ninth before or after the rest is the order that
4888 // could be wrong, so it is refused instead.
4889 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4890 .expect_err("nine do not fit on the stack");
4891 assert_eq!(
4892 failed.to_string(),
4893 "block1 takes 9 parameters of type `f80` and only 8 can cross an edge at once"
4894 );
4895 assert_eq!(failed.inst(), None);
4896 }
4897}