rucc_codegen/lower.rs
1//! The selector: an IR function becomes a machine IR function.
2//!
3//! Design: `spec/10-backend.md` sections 10.2 and 10.3.
4//!
5//! What the matcher in [`crate::select`] does is answer one question about one term. What this
6//! does is ask it: walk a function, decide which terms are worth asking about, and build machine
7//! instructions out of what comes back. Nothing here decides what an IR term lowers to. That is
8//! in `rules/x86-64.rules` and it is proved before it is used, which is the whole point of the
9//! arrangement and the reason this file is short.
10//!
11//! # What it does with an instruction
12//!
13//! It tries the ways the instruction can be shown to the matcher, in order, and takes the first
14//! that a rule fires on. [`crate::term`] is what a way of showing one is, and the order is the
15//! most specific first: an operand that is a constant is offered as a constant before it is
16//! offered as a register, and an operand computed by an instruction of its own is offered as
17//! that instruction before it is offered as a register. A rule that wants an immediate too wide
18//! for the machine has a guard that turns it down, and the search carries on to the way of
19//! showing it that puts the constant in a register, which is the right answer and is one nobody
20//! had to write down.
21//!
22//! A constant is not lowered where it is written. It is materialized where a register for it is
23//! first wanted, which is what keeps a constant that every use folded into an immediate from
24//! leaving a dead instruction behind, and it also gives the value the shortest live range it
25//! could have. The instruction that materializes it comes from the rule set like everything else.
26//!
27//! # What it does not do yet
28//!
29//! Everything is in the general purpose registers, because every rule in the set is about an
30//! integer, so a call that passes a `double` and a function that returns one are both reported
31//! rather than lowered. So is an argument that travels on the stack, on either side of a call,
32//! and so is a call through an address rather than to a name.
33//!
34//! # A call
35//!
36//! Not a rule, because a rule pattern sees one term and what a call's operands are is whatever
37//! the signature made them. [`crate::abi`] builds one instead, out of the same description of the
38//! convention the arguments come from: the values it passes are reads constrained to the
39//! registers the convention places them in, what comes back is a write constrained to the
40//! register it comes back in, and every other register the callee is free to destroy is a write
41//! of that register and nothing else, which is all the allocator needs to keep a value out of it.
42//!
43//! What that costs the frame is an argument area, and nothing after selection could work out how
44//! big, so the size of the widest call is given back with the function. A function that makes no
45//! call at all is a leaf, and a leaf is the function that may use the red zone.
46//!
47//! # Where a block goes
48//!
49//! On the block, which is what machine IR does with an edge and is why the branches need no more
50//! rule language than the arithmetic did. A rule never names a block, so an unconditional jump
51//! has no rule at all and a conditional branch has one that is about its condition and nothing
52//! else. The arms are copied across after the block is filled, arguments and all, because an
53//! argument that is a constant is materialized where a register for it is first wanted and the
54//! end of the block is where an edge wants it.
55//!
56//! What this leaves behind is a function whose blocks are in the order the IR held them and whose
57//! branches are still branches on a register. Turning one into a `test` and a `jcc` is the block
58//! layout's, since which of the two arms falls through is the layout's answer, and [`crate::split`]
59//! has to run before allocation so that every edge carrying a value has somewhere to put it.
60//!
61//! A store and a return are the two things here that write no register. A store is emitted like
62//! everything else and the only difference is that there is no result to put anywhere, so the
63//! operands the target describes are all reads. A return is the same, and what it is for is its
64//! one operand: the target constrains it to the register the caller reads the value out of, and
65//! the allocator is what gets it there. The instruction that leaves is not chosen here at all,
66//! because the epilogue has to give the frame back first and [`crate::finish`] writes that after
67//! allocation, so a return of nothing is lowered to nothing.
68//!
69//! The entry block is the one block whose parameters are not block parameters here. They are the
70//! function's arguments, they are already somewhere when it starts, and [`crate::abi`] is what
71//! says where. An argument that arrives on the stack is reported rather than read, because where
72//! the stack put it is a distance into a frame and no frame exists until after allocation.
73//!
74//! Blocks are walked in the order the function holds them and a value is expected to be defined
75//! before it is used, which is true of the IR this is given because every pass before it keeps
76//! definitions ahead of uses.
77
78use std::collections::HashSet;
79use std::fmt;
80
81use rucc_base::{Interner, Symbol};
82use rucc_diag::Span;
83use rucc_ir::{
84 Abi, AsmOperand, AsmOperands, Block, Def, Extra, FloatPred, Func, Inst, Linkage, MemOrder,
85 Opcode, Param, PrefetchHint, RmwOp, Type, Value, Visibility,
86};
87use rucc_mir as mir;
88use rucc_target::x86_64;
89use rucc_target::{CallRegs, Constraint, OperandDesc, PhysReg, RegClass, Role, Segment};
90
91use crate::abi::{self, Missing, Refused};
92use crate::coverage::Fired;
93use crate::elsewhere::Elsewhere;
94use crate::frame::{Layout, Local};
95use crate::select::{Match, Piece, Rule, Table};
96use crate::term::{MAX_ARGS, PLAIN, Plan, Shown, Term, Terms};
97use crate::varargs;
98
99/// The prefix a rule file puts in front of a machine term, which says which target it belongs
100/// to and is not part of the opcode.
101pub(crate) const PREFIX: &str = "x64.";
102
103/// The instruction a global offset table slot is read with.
104///
105/// Not in [`x86_64::FRAME`] with the other opcodes this file names, because a frame has no use for
106/// it. It is spelled out here because the relocation it takes is only legal on a `mov` with a REX
107/// prefix, so the width is part of the requirement rather than a choice.
108const GOT_LOAD: &str = "mov_rm_64";
109
110/// How wide an address is on this target, which is the width a cast between a pointer and an
111/// integer has to be at for the cast to be nothing.
112const ADDRESS_BITS: u32 = 64;
113
114/// How many bytes a `long double` takes in memory, and what it is aligned to, which are the same
115/// number and are both more than the ten bytes that mean anything.
116///
117/// The psABI's answer rather than a choice here. `sizeof (long double)` is sixteen on this
118/// machine, so an array of them is laid out this way whatever a slot holding one does, and a slot
119/// that agreed with the array is one fewer thing to get wrong.
120const X87_BYTES: u32 = 16;
121
122/// How many values the x87 stack holds at once.
123///
124/// Eight, which is the machine's number rather than a choice here, and it matters in one place:
125/// the parameters of a block are copied through the stack so that they all move at once, and a
126/// block with more of them than this has nowhere to put the ninth.
127const X87_DEPTH: usize = 8;
128
129/// How many bytes a value passes through on its way between a register and the x87 stack.
130///
131/// Eight, because the widest thing that crosses is a `double` or a sixty four bit integer, and
132/// nothing crosses at eighty bits: a value that wide is already in the frame and the stack reaches
133/// it where it is.
134const X87_CROSSING: u32 = 8;
135
136/// Where the rounding field of the x87 control word is and what it has to be set to for the unit
137/// to cut towards zero, which is the one rounding C asks for that the unit does not do by default.
138///
139/// Both bits on is truncate. The field is ORed into the word that was already there rather than
140/// written over it, so the precision control and the exception masks somebody else set stay set.
141const X87_TRUNCATE: i64 = 0x0c00;
142
143/// Whether a type is the one this machine has no register for.
144///
145/// Only the eighty bit float is, and that is a fact about x86-64 rather than about floats: every
146/// other scalar the front end produces is in a general purpose register or a vector one, and this
147/// one is on the x87 stack while it is being worked on and in memory the rest of the time. So it
148/// has no place in [`Lowering::class_of`] and no name in [`crate::term`], and every instruction
149/// that touches one is written out by hand in this file.
150fn on_x87(ty: Type) -> bool {
151 ty.is_scalar() && ty.is_float() && ty.bits() == 80
152}
153
154/// Where one operand of an assembly statement is, on each side of the assembly.
155///
156/// Two registers rather than one, because an operand written `+` is a value that arrives and a
157/// value that leaves and those are two values. The machine IR has one definition per register by
158/// construction, so an instruction of the template that reads the operand and writes it has to name
159/// a different register in each place, and what makes the two one register in the end is the
160/// [`Constraint::Reuse`] the instruction's description carries: the allocator reads it, gives both
161/// the same physical register, and copies the incoming value somewhere first when something else is
162/// still using it.
163///
164/// Most operands have one of the two. An input has only a place it is read from and an output
165/// written `=` has only a place it is written to, and asking either of them for the other is an
166/// operand read where the opcode writes or written where it reads, which [`Lowering::placed`]
167/// refuses.
168#[derive(Debug, Clone, Copy, Default, PartialEq, Eq)]
169struct Place {
170 /// The register the value arrives in, for an operand something reads.
171 read: Option<mir::Reg>,
172 /// The register the value leaves in, for an operand something writes.
173 write: Option<mir::Reg>,
174}
175
176/// Which of an assembly statement's operands is in that register, for an instruction that reaches
177/// the register without its text saying so.
178///
179/// The constraint letter is what says so, and it is the only thing in such a statement that could:
180/// `"=a"` is an output in `rax` and `"c"` is an input in `rcx`, and a register nothing names is a
181/// register nobody has said anything about. So a write looks among the outputs and a read among the
182/// inputs, and an output written `+` answers for either, since it is read before it is written.
183///
184/// `None` is a register the instruction uses and the statement put nothing in, which is the usual
185/// answer rather than an unusual one. `cpuid` writes four registers and a program that wanted one
186/// of them names one. See [`Lowering::spare`], which is where that one goes.
187fn bound(list: &[AsmOperand], reg: PhysReg, role: Role) -> Option<usize> {
188 list.iter().position(|operand| {
189 operand.fixed.and_then(x86_64::gpr_letter) == Some(reg)
190 && if role.is_def() { operand.result.is_some() } else { operand.value.is_some() }
191 })
192}
193
194/// Why a function could not be lowered.
195///
196/// One reason and then nothing. A function with no rule for something in it is a function this
197/// cannot finish, and the second thing it could not lower is not news.
198#[derive(Debug, Clone, PartialEq, Eq)]
199pub enum Unsupported {
200 /// An instruction no rule fires on.
201 Inst {
202 /// The instruction that stopped it.
203 inst: Inst,
204 /// What the rule file would call it, or nothing if the rule language has no name for it
205 /// at all, which is what an instruction at a width nothing is written about looks like.
206 term: Option<&'static str>,
207 /// The opcode, which is what gets named when the rule language has no word for it.
208 ///
209 /// An opcode the rule language has no word for is exactly the opcode no rule lowers, so
210 /// without this the message would be empty in every case where somebody needs it.
211 opcode: Opcode,
212 /// What it produces, or nothing for an instruction that is only an effect.
213 ty: Option<Type>,
214 },
215 /// A parameter that does not arrive somewhere this can bring it in from.
216 ///
217 /// Not an instruction, which is why it is a separate arm: it is a fact about the signature
218 /// and there is nothing in the body of the function to point at.
219 Argument {
220 /// Its position in the signature.
221 index: usize,
222 /// What is wrong with where it arrives.
223 missing: Missing,
224 },
225 /// A call that passes or gives back a value this cannot put where the convention wants it.
226 Call {
227 /// The call.
228 inst: Inst,
229 /// Which value, and what is wrong with where it travels.
230 refused: Refused,
231 },
232 /// A `return` this cannot put where the convention wants it.
233 ///
234 /// A separate arm from [`Unsupported::Inst`] because it is not an instruction no rule fires
235 /// on. A return of more than one value is built from the convention rather than matched, the
236 /// same way a call is, so what goes wrong with one is what goes wrong with a call and not the
237 /// absence of a rule.
238 Returned {
239 /// The `return`.
240 inst: Inst,
241 /// What is wrong with where one of the values travels.
242 missing: Missing,
243 },
244 /// A stack slot the frame cannot give the bytes it asked for.
245 ///
246 /// Not an instruction no rule covers. An `alloca` is built here rather than matched, so what
247 /// goes wrong with one is what the frame can and cannot hold rather than what the rules spell.
248 Dynamic {
249 /// The `alloca`.
250 inst: Inst,
251 /// What the frame could not do about it.
252 growing: Growing,
253 },
254 /// More parameters of a type that travels on the x87 stack than the stack is deep.
255 ///
256 /// Not an instruction either, for the reason a function's parameter is not one: it is a fact
257 /// about the block and there is nothing in the block to point at. What crosses an edge for one
258 /// of these is the address of where the value is, and the block copies the bytes into a slot
259 /// of its own, all of them through the stack at once so that a block carrying two of them
260 /// swapped is copied in an order that is right. Eight is as many as the stack holds, and a
261 /// ninth would have to be copied before or after the rest, which is the order that could be
262 /// wrong.
263 Phi {
264 /// Which block it arrives at.
265 block: Block,
266 /// How many of them arrive there, which is the whole of what is wrong.
267 count: usize,
268 /// What they are.
269 ty: Type,
270 },
271 /// An `asm` statement this cannot build.
272 ///
273 /// Not an instruction no rule fires on, for the reason a call is not one: what it stands for is
274 /// whatever its template says, and no pattern over terms can read a string.
275 Assembly {
276 /// The `inline_asm`.
277 inst: Inst,
278 /// What about it is not built here yet.
279 refused: Written,
280 },
281}
282
283/// What about an `asm` statement is not built yet.
284#[derive(Debug, Clone, Copy, PartialEq, Eq)]
285pub enum Written {
286 /// A template with instructions in it.
287 Template,
288 /// An `asm goto`, whose labels make the statement a terminator.
289 Goto,
290 /// An operand this cannot put where the constraint says it goes.
291 Operand,
292 /// A clobber list naming something this has no register for.
293 Clobber,
294}
295
296impl Written {
297 /// The rest of the sentence that starts with the statement.
298 #[must_use]
299 pub fn why(self) -> &'static str {
300 match self {
301 // The template is the assembler's to read and there is no assembler here yet, so a
302 // template with anything in it is a string nothing can turn into bytes. An empty one is
303 // no instructions, and no instructions is something this can write.
304 Written::Template => "has instructions in its template, which nothing here assembles",
305 Written::Goto => "jumps to a label, which nothing here builds an edge for",
306 Written::Operand => "has an operand this cannot place",
307 Written::Clobber => "says it destroys a register this has no name for",
308 }
309 }
310}
311
312/// What the frame could not do about a stack slot.
313#[derive(Debug, Clone, Copy, PartialEq, Eq)]
314pub enum Growing {
315 /// An object of a size the number a frame counts bytes in does not reach.
316 Huge,
317 /// A variable length array wanting more alignment than a call leaves the stack pointer with.
318 ///
319 /// Rounding the stack pointer down again after the bytes have been taken would put it
320 /// somewhere no constant reaches the rest of the frame from, so a frame like this needs a
321 /// second base register held for the whole of the function. Nothing here holds one.
322 Aligned,
323}
324
325impl Growing {
326 /// The rest of the sentence that starts with the slot.
327 #[must_use]
328 pub fn why(self) -> &'static str {
329 match self {
330 Growing::Huge => "is more bytes than a frame counts",
331 Growing::Aligned => {
332 "wants more alignment than the stack pointer is left on, which needs a base \
333 register nothing here keeps"
334 }
335 }
336 }
337}
338
339impl Unsupported {
340 /// The instruction it is about, or nothing for the one arm that is about a signature.
341 ///
342 /// What a caller wants this for is the span. The function knows where every instruction in
343 /// it came from, so a caller holding both can point a message at the line somebody wrote
344 /// rather than at the file as a whole, and nothing here has to carry a span of its own.
345 pub fn inst(&self) -> Option<Inst> {
346 match *self {
347 Unsupported::Inst { inst, .. }
348 | Unsupported::Call { inst, .. }
349 | Unsupported::Returned { inst, .. }
350 | Unsupported::Dynamic { inst, .. }
351 | Unsupported::Assembly { inst, .. } => Some(inst),
352 Unsupported::Argument { .. } | Unsupported::Phi { .. } => None,
353 }
354 }
355}
356
357impl fmt::Display for Unsupported {
358 fn fmt(&self, f: &mut fmt::Formatter<'_>) -> fmt::Result {
359 match *self {
360 Unsupported::Inst { term: Some(term), .. } => write!(f, "no rule lowers `{term}`"),
361 Unsupported::Inst { term: None, opcode, ty: Some(ty), .. } => {
362 write!(f, "no rule lowers a `{opcode}` producing a `{ty}`")
363 }
364 Unsupported::Inst { term: None, opcode, ty: None, .. } => {
365 write!(f, "no rule lowers a `{opcode}`")
366 }
367 Unsupported::Argument { index, missing } => {
368 write!(f, "parameter {index} {}", missing.why())
369 }
370 Unsupported::Call { refused: Refused { argument: Some(index), missing }, .. } => {
371 write!(f, "argument {index} of this call {}", missing.why())
372 }
373 Unsupported::Call { refused: Refused { argument: None, missing }, .. } => {
374 write!(f, "what this call gives back {}", missing.why())
375 }
376 Unsupported::Returned { missing, .. } => {
377 write!(f, "what this function gives back {}", missing.why())
378 }
379 Unsupported::Dynamic { growing, .. } => {
380 write!(f, "this local {}", growing.why())
381 }
382 Unsupported::Phi { block, count, ty } => {
383 let block = block.index();
384 write!(
385 f,
386 "block{block} takes {count} parameters of type `{ty}` and only {X87_DEPTH} can cross an edge at once"
387 )
388 }
389 Unsupported::Assembly { refused, .. } => write!(f, "this `asm` {}", refused.why()),
390 }
391 }
392}
393
394impl std::error::Error for Unsupported {}
395
396/// A lowered function, and what the frame needs that the machine IR does not hold.
397#[derive(Debug)]
398pub struct Lowered {
399 /// The function, in machine instructions.
400 pub func: mir::Func,
401 /// What it wants its stack to look like, which is separate from the function so that the two
402 /// can be read and written at the same time.
403 pub stack: Stack,
404 /// Which rules of the table lowered it, which is what `-Zrule-coverage` asks for and what
405 /// `crate::coverage` writes down.
406 pub fired: Fired,
407 /// Which machine IR block each IR block became, indexed by the IR block's own index, and
408 /// nothing for a block the walk never reached.
409 ///
410 /// Here because it is the only place the correspondence exists. Selection makes one block per
411 /// block, in the same order and with the arms in the same order, so anything the IR knows
412 /// about a block can be carried down through this and nothing else, and
413 /// [`crate::weights::carry`] is what does.
414 pub blocks: Vec<Option<mir::Block>>,
415}
416
417/// What a function's stack has to hold, as far as selection is able to say.
418///
419/// All of it is answered here because selection is where a call is built and where an `alloca`
420/// is read, and nothing after it could tell what either of them needed.
421#[derive(Debug, Default)]
422pub struct Stack {
423 /// How many bytes the widest call in the function needs below the stack pointer for the
424 /// arguments it passes there, or `None` for a function that makes no call at all.
425 ///
426 /// `None` is a leaf, which is the function that may use the red zone and the one whose stack
427 /// pointer does not have to be left aligned for anybody.
428 pub calls: Option<u32>,
429 /// The memory the function asked for itself, one entry for every `alloca` in it, in the order
430 /// the walk reached them.
431 pub locals: Vec<Local>,
432 /// Which instruction computes the address of which of those locals.
433 ///
434 /// An address in the frame is a distance from the stack pointer, and there is no frame until
435 /// after allocation, so the instruction is written here with nothing in its displacement and
436 /// [`crate::finish`] writes the number in once [`crate::frame::Frame`] knows it.
437 pub addresses: Vec<(mir::Inst, usize)>,
438 /// Which instruction computes the address of a piece of memory whose size the function works
439 /// out while it runs, which is what a variable length array is.
440 ///
441 /// Waiting on [`crate::finish`] for a different number from the one the addresses above are:
442 /// the bytes were taken off the stack pointer by the instruction in front of this one, so where
443 /// they start is however much of the bottom of the frame belongs to the arguments of a call,
444 /// and that is not known until the frame is.
445 pub dynamic: Vec<mir::Inst>,
446 /// Which instruction takes those bytes off the stack pointer, one for every one of them, in the
447 /// order the walk reached them.
448 ///
449 /// Read by [`crate::finish`] on a command line that asked for the stack to be touched a page at
450 /// a time, which is the one thing that has to find these again: the bytes are in a register by
451 /// then, so the walk down to them is a loop, and a loop is written around an instruction rather
452 /// than in front of a block. Nothing else looks at them, because everything else about a frame
453 /// that grows is answered by the address the instruction below this one computes.
454 pub grown: Vec<mir::Inst>,
455 /// Where the function first moves the stack pointer while it runs, if it does at all.
456 ///
457 /// Two things are read off this. One is whether at all, which is what [`crate::frame::Layout`]
458 /// wants, because a frame that moves its stack pointer has a different shape from one that does
459 /// not and the layout is built before the instructions are looked at again. See `Growing` in
460 /// [`crate::frame`]. The other is where, so that a caller that cannot accept such a frame has
461 /// somewhere to point when it says so.
462 pub grown_at: Option<Inst>,
463 /// Which instruction reads which of the arguments the caller passed on the stack, as how far up
464 /// the caller's argument area it reads.
465 ///
466 /// Waiting on [`crate::finish`] for the same reason the addresses above are, and on one thing
467 /// more: where the caller's argument area is from inside this function depends on whether the
468 /// prologue had to force the stack pointer's alignment, so which register the load reads
469 /// through is not settled here either.
470 pub arguments: Vec<(mir::Inst, u32)>,
471 /// Whether the function asked where its own frame is, which is what `__builtin_frame_address`
472 /// and `__builtin_return_address` both start from.
473 ///
474 /// A function like that keeps a frame pointer whatever the flags say, because the register is
475 /// the answer to the first of them and the start of the walk for every depth above zero. There
476 /// is no other way to reach it: the distance from the stack pointer to the frame is a number
477 /// the layout works out, and what a walk up the chain needs is the link the prologue saved.
478 pub walks_frames: bool,
479}
480
481impl Stack {
482 /// The layout given, with the three fields only the lowering knows the answer to filled in.
483 ///
484 /// Everything else in a layout comes from the flags the function is compiled under or from the
485 /// allocation, so this takes one and returns it rather than building one.
486 #[must_use]
487 pub fn layout<'a>(&'a self, base: Layout<'a>) -> Layout<'a> {
488 Layout {
489 leaf: self.calls.is_none(),
490 outgoing: self.calls.unwrap_or(0),
491 locals: &self.locals,
492 grows: self.grown_at.is_some(),
493 ..base
494 }
495 }
496}
497
498/// The x86-64 machine IR for that function.
499///
500/// # Errors
501///
502/// The first instruction no rule fires on, which today is anything at a width the rule set is not
503/// written at, a parameter that does not arrive in a register this can read, or a call that
504/// passes something this cannot put where the convention wants it.
505pub fn func(
506 source: &Func,
507 names: &mut Interner,
508 conv: &'static CallRegs,
509 elsewhere: &Elsewhere,
510) -> Result<Lowered, Unsupported> {
511 Lowering::new(source, names, conv, elsewhere).run()
512}
513
514/// What the matcher settled on for one block, indexed the way the block's instructions are.
515struct Decided {
516 /// What each instruction matched, and nothing for one that matched no rule or was folded
517 /// into a later one.
518 found: Vec<Option<Match<Term>>>,
519 /// How each instruction showed its operands to the matcher, which is what says what it took.
520 plans: Vec<Option<Plan>>,
521 /// The instructions some other instruction took, which are the ones with nothing to write.
522 folded: Vec<Inst>,
523}
524
525/// One function being lowered.
526struct Lowering<'a> {
527 source: &'a Func,
528 names: &'a mut Interner,
529 out: mir::Func,
530 /// The machine register each IR value is in, once it has one.
531 regs: Vec<Option<mir::Reg>>,
532 /// For a constant that has been written into a register, the block it was written into,
533 /// which is the only block that register is any good in.
534 written: Vec<Option<mir::Block>>,
535 /// How many times each IR value is read, which is what says whether an instruction may be
536 /// folded into the one that reads it.
537 uses: Vec<u32>,
538 /// The block being filled.
539 at: Option<mir::Block>,
540 /// The machine IR block each IR block became.
541 blocks: Vec<Option<mir::Block>>,
542 /// The class an address is in, which is the general purpose one and is not a question: every
543 /// register an addressing mode names holds part of an address, and there is no machine here
544 /// that computes an address anywhere but in this file. Which class a *value* is in is
545 /// [`Lowering::class_of`], and it is a question, because a float is in the other one.
546 gpr: RegClass,
547 /// Where the convention this function is compiled for puts things, which is read for the
548 /// arguments and for the calls.
549 conv: &'static CallRegs,
550 /// Which names this function may not work an address out for itself, which is a fact about the
551 /// module and so is worked out before any of this and handed in.
552 elsewhere: &'a Elsewhere,
553 /// What the function wants its stack to look like, filled in as the walk finds out.
554 stack: Stack,
555 /// What a `va_start` in this function has to write, or nothing for a function that takes no
556 /// arguments its signature does not name.
557 ///
558 /// Worked out once, when the entry block binds the parameters, because every number in it is
559 /// about where those parameters left the walk over the argument registers and there is nowhere
560 /// else that knows.
561 varargs: Option<Varargs>,
562 /// Which of the function's stack objects each eighty bit value lives in, once it has asked
563 /// for one.
564 ///
565 /// One slot per value and it is never given back, which is what makes an eighty bit value
566 /// behave like every other one: it is written once and read wherever it is read, and no two
567 /// of them share a slot the way two of them would share a register. What is in a register is
568 /// the address, and that is worked out again at every use rather than kept, so nothing here
569 /// holds a general purpose register open across a whole function.
570 slots: Vec<Option<usize>>,
571 /// The eight bytes a value passes through between a register and the x87 stack, once
572 /// something has wanted them.
573 ///
574 /// One for the whole function, because every group that uses it is a handful of instructions
575 /// with nothing in between: the bytes are written, read straight back and never looked at
576 /// again, so a second slot would be a second slot holding the same nothing.
577 crossing: Option<usize>,
578 /// The four bytes the control word is saved in and the changed copy written to, once
579 /// something has wanted them.
580 ///
581 /// One for the whole function for the reason above, and four rather than two because it is
582 /// two words: the one the unit had and the one with the rounding field turned to truncate.
583 control: Option<usize>,
584 /// Which rules have fired so far.
585 fired: Fired,
586}
587
588/// What a `va_start` in a variadic function writes into the list it is given.
589///
590/// Three of the four are settled here and the fourth is not a number at all yet: where the save
591/// area is and where the caller's argument area is are both distances into a frame that does not
592/// exist until after allocation, so both are `lea` instructions [`crate::finish`] fills in.
593#[derive(Debug, Clone, Copy, PartialEq, Eq)]
594struct Varargs {
595 /// Which of the function's stack objects is the register save area.
596 save: usize,
597 /// How far up the caller's argument area the first argument the signature does not name is,
598 /// which is the whole of that area the named ones did not take.
599 incoming: u32,
600 /// What `gp_offset` starts at, which is past the general purpose registers the named arguments
601 /// took.
602 integers: u32,
603 /// What `fp_offset` starts at, which is past the vector ones.
604 floats: u32,
605}
606
607/// How far a function's name reaches, narrowed from the linkage the IR gave it.
608///
609/// The IR has five and an object file says three, and the two the linker cannot tell apart are
610/// the two weak ones: which of them a symbol had is a fact the optimizer reads and the linker has
611/// no way to record. A function is never `Common`, since that is what a tentative definition of an
612/// object is and there is no tentative definition of a function, and it is written here rather
613/// than left out so that a linkage added later has to come past this.
614const fn binding(linkage: Linkage) -> mir::Binding {
615 match linkage {
616 Linkage::Internal => mir::Binding::Local,
617 Linkage::Weak | Linkage::LinkOnce => mir::Binding::Weak,
618 Linkage::External | Linkage::Common => mir::Binding::Global,
619 }
620}
621
622/// How far a function's name reaches outside a shared library, carried across unchanged.
623///
624/// Nothing is narrowed here the way [`binding`] narrows the linkage, because ELF records all
625/// three of these and the two enumerations are the same three answers written twice: once in a
626/// crate that is not allowed to know what an object file is and once in one that is.
627const fn visibility(visibility: Visibility) -> mir::Visibility {
628 match visibility {
629 Visibility::Default => mir::Visibility::Default,
630 Visibility::Hidden => mir::Visibility::Hidden,
631 Visibility::Protected => mir::Visibility::Protected,
632 }
633}
634
635impl<'a> Lowering<'a> {
636 fn new(
637 source: &'a Func,
638 names: &'a mut Interner,
639 conv: &'static CallRegs,
640 elsewhere: &'a Elsewhere,
641 ) -> Self {
642 let counts = source.counts();
643 let name = source.name;
644 let mut uses = vec![0; counts.values];
645 for block in source.blocks() {
646 for inst in source.insts(block) {
647 for &arg in &source[source[inst].args] {
648 uses[arg.index()] += 1;
649 }
650 for call in source.successors(inst) {
651 for &arg in &source[call.args] {
652 uses[arg.index()] += 1;
653 }
654 }
655 }
656 }
657 let mut out = mir::Func::new(name);
658 out.align = source.align;
659 out.binding = binding(source.linkage);
660 out.visibility = visibility(source.visibility);
661 Self {
662 source,
663 names,
664 out,
665 regs: vec![None; counts.values],
666 written: vec![None; counts.values],
667 blocks: vec![None; counts.blocks],
668 uses,
669 at: None,
670 gpr: x86_64::GPR,
671 conv,
672 elsewhere,
673 stack: Stack::default(),
674 varargs: None,
675 slots: vec![None; counts.values],
676 crossing: None,
677 control: None,
678 fired: Fired::new(),
679 }
680 }
681
682 fn run(mut self) -> Result<Lowered, Unsupported> {
683 // Every block before any of them is filled, because a block that jumps forward has to
684 // name the block it jumps to and a machine IR block is named by a handle rather than by
685 // the IR block it came from.
686 for block in self.source.blocks() {
687 let out = self.out.create_block();
688 self.blocks[block.index()] = Some(out);
689 }
690 for block in self.order() {
691 self.block(block)?;
692 }
693 // And the name each block an image holds the address of was given, which nothing in the
694 // walk above would ask for: the `lea` a label address is inside the function needs no
695 // symbol, and the one thing that does is a relocation in another section.
696 let named: Vec<(Block, Symbol)> = self.source.named_blocks().collect();
697 let labels: Vec<(mir::Block, Symbol)> =
698 named.into_iter().map(|(block, name)| (self.out_block(block), name)).collect();
699 self.out.labels = labels;
700 Ok(Lowered { func: self.out, stack: self.stack, fired: self.fired, blocks: self.blocks })
701 }
702
703 /// The order the blocks are filled in, which is not the order they are written in.
704 ///
705 /// Reverse postorder, because a value is written in a block that dominates every block that
706 /// reads it and a block in reverse postorder comes before every block it dominates. The order
707 /// the blocks are written in does not have that property: a block written early can read a
708 /// value a block below it writes, and reading a value with no register yet mints one, so the
709 /// register the definition writes later is not the register the read named. Nothing writes the
710 /// one the read named, and what comes out is a function that loads a stack slot no store ever
711 /// reached. It is the order this walk goes in rather than the order the blocks come out in,
712 /// which is what the loop above fixes, so the machine function is still written the way the IR
713 /// function was.
714 ///
715 /// Blocks the entry does not reach come last, in the order they are written in. Nothing runs
716 /// them and nothing they name is read by anything that does, but they still have to be filled,
717 /// because a machine block with no terminator is not one the passes below can read.
718 fn order(&self) -> Vec<Block> {
719 let Some(entry) = self.source.entry() else { return self.source.blocks().collect() };
720 let count = self.blocks.len();
721 let mut succs: Vec<Vec<Block>> = vec![Vec::new(); count];
722 for block in self.source.blocks() {
723 let Some(term) = self.source.terminator(block) else { continue };
724 succs[block.index()] = self.source.successors(term).map(|call| call.block).collect();
725 }
726 // An explicit stack, because the depth of the walk is the number of blocks and a function
727 // built by a generator has as many of those as it likes.
728 let mut seen = vec![false; count];
729 let mut order = Vec::with_capacity(count);
730 let mut stack = vec![(entry, 0usize)];
731 seen[entry.index()] = true;
732 while let Some((block, at)) = stack.pop() {
733 let Some(&next) = succs[block.index()].get(at) else {
734 order.push(block);
735 continue;
736 };
737 stack.push((block, at + 1));
738 if !seen[next.index()] {
739 seen[next.index()] = true;
740 stack.push((next, 0));
741 }
742 }
743 order.reverse();
744 order.extend(self.source.blocks().filter(|block| !seen[block.index()]));
745 order
746 }
747
748 /// One block: its parameters, then every instruction in it that is not folded into another.
749 fn block(&mut self, block: Block) -> Result<(), Unsupported> {
750 let out = self.out_block(block);
751 self.at = Some(out);
752 if self.source.entry() == Some(block) {
753 self.arrive(block, out)?;
754 } else {
755 let mut arriving = Vec::new();
756 for ¶m in &self.source[block].params {
757 // A value with no register to arrive in, which the class would not say, since
758 // `class_of` puts one of these in the general purpose file on purpose and what it
759 // means by that is that nothing there can hold it. What crosses the edge for one
760 // of those is the address of where the value already is, so the parameter is a
761 // pointer here and the bytes it points at are copied below.
762 let ty = self.source[param].ty;
763 let reg = self.out.append_param(out, self.class_of(ty));
764 self.regs[param.index()] = Some(reg);
765 if on_x87(ty) {
766 arriving.push((param, reg));
767 }
768 }
769 self.settle(block, &arriving)?;
770 }
771
772 // What each instruction matched, and which instructions were folded into another. The
773 // decision is made for the whole block before any of it is written, and it is made more
774 // than once: a value that only some of its readers took has to be put back in a register
775 // for all of them, and taking it away from those readers changes what they match.
776 let insts: Vec<Inst> = self.source.insts(block).collect();
777 let mut refused: HashSet<Value> = HashSet::new();
778 let mut decided = self.decide(&insts, &refused);
779 while let Some(value) = self.left_alive(&insts, &decided.plans) {
780 refused.insert(value);
781 decided = self.decide(&insts, &refused);
782 }
783 let Decided { found, folded, .. } = decided;
784
785 for (&inst, matched) in insts.iter().zip(found) {
786 if folded.contains(&inst) || self.writes_nothing(inst) {
787 continue;
788 }
789 // A call is built from the convention rather than matched, which is why it is the one
790 // opcode looked at by name here. Through an address it is a different instruction and
791 // the same convention, so the two arrive at the same place and differ in one line of
792 // it.
793 match self.source[inst].opcode {
794 Opcode::Call | Opcode::CallIndirect => {
795 self.called(inst)?;
796 continue;
797 }
798 // Built from the frame rather than matched, for the same shape of reason a call
799 // is built from the convention: what a rule replaces a term with is instructions,
800 // and what an `alloca` needs first is bytes, which the rule language has no way
801 // to ask for.
802 Opcode::Alloca => {
803 self.reserve(inst)?;
804 continue;
805 }
806 // Reading the stack pointer and writing it back, which are the two ends of a scope
807 // holding a variable length array. Built here for the reason an `alloca` is: the
808 // value is a register the rule language has no way to name, because what it holds
809 // is not a value the program computed but where the machine's stack had got to.
810 Opcode::StackSave => {
811 self.stack_pointer(inst, false)?;
812 continue;
813 }
814 Opcode::StackRestore => {
815 self.stack_pointer(inst, true)?;
816 continue;
817 }
818 // The address of a name, built here for the same reason an `alloca` is: what a
819 // rule replaces a term with is instructions over values, and the operand of this
820 // one is a symbol, which is a thing the rule language has no way to bind and the
821 // solver has no way to say anything about. There is nothing in `lea sym(%rip)` a
822 // proof over bitvectors could discharge, because what makes it the right answer
823 // is the relocation and what the linker does with it.
824 Opcode::GlobalAddr => {
825 self.address_of(inst)?;
826 continue;
827 }
828 // The address of a label and the branch that reads one, built here for the same
829 // reason and for one more. The reason is the same: what the first of them names is
830 // a block, which is not a value a rule pattern can bind, and there is nothing in
831 // the distance between two places in one function that a proof over bitvectors
832 // could discharge. The extra one is that the second is a terminator whose arms are
833 // not two and not fixed, and a rule says what an instruction reads rather than
834 // where a block goes.
835 Opcode::BlockAddr => {
836 self.block_address(inst)?;
837 continue;
838 }
839 Opcode::IndirectBr => {
840 self.indirect_branch(inst)?;
841 continue;
842 }
843 // Where this thread's own storage starts, built here for a reason of the same
844 // shape: what it reads is `%fs`, which is not a register the rule language can
845 // bind and not one a proof over bitvectors could say anything about, because what
846 // makes the load the right answer is an agreement between the loader and the C
847 // library rather than any arithmetic.
848 Opcode::ThreadPointer => {
849 self.thread_pointer(inst)?;
850 continue;
851 }
852 // Where a frame is and what it returns to, built here for the same reason and one
853 // more. The reason is the same: what the walk starts from is the frame pointer,
854 // which is not a register a rule pattern can bind, and there is nothing in reading
855 // the link the prologue saved that a proof over bitvectors could discharge. The
856 // extra one is that how long the walk is comes out of a number beside the
857 // instruction, so one of these is not one instruction but however many the depth
858 // says, and a rule replaces a term with a term.
859 Opcode::FrameAddress | Opcode::ReturnAddress => {
860 self.frames(inst)?;
861 continue;
862 }
863 // Built from the frame for the reason an `alloca` is, and from the convention for
864 // the reason a call is: three of the four fields it writes are distances that do
865 // not exist until the frame does, and the fourth is where the walk over the
866 // argument registers stopped. A function that is not variadic has no such walk to
867 // report, so it has nothing here and is refused below, which is the right answer
868 // for a `va_start` in one.
869 Opcode::VaStart if self.varargs.is_some() => {
870 self.va_start(inst)?;
871 continue;
872 }
873 // A return of more than one value, which is a structure small enough to come
874 // back in a pair of registers. Built from the convention for the reason a call
875 // is: which register each half goes in depends on the halves in front of it,
876 // because the two register files are walked separately, and a pattern over a term
877 // cannot see them. A return of one value is a term with a name and a rule, and it
878 // stays one.
879 //
880 // A return of none in a function whose answer went through memory is here too,
881 // and for a different reason: what it gives back is not written in the IR at all.
882 // The convention says the address the caller handed over comes back, and only the
883 // signature says this function was handed one.
884 //
885 // And a return of one eighty bit value, for a third reason: what a rule would
886 // write is an instruction leaving the value in a register, and this one is left on
887 // the x87 stack instead. A rule could not name that stack any more than any other
888 // rule about this type could.
889 Opcode::Return
890 if self.source[self.source[inst].args].len() > 1
891 || self.sret().is_some()
892 || self.gives_back_x87(inst) =>
893 {
894 self.returned(inst)?;
895 continue;
896 }
897 // A cast between a pointer and an integer of the same width, which on this
898 // machine is every one the front end writes. No instruction at all, so no rule
899 // could name one.
900 Opcode::PtrToInt | Opcode::IntToPtr => {
901 self.rename(inst)?;
902 continue;
903 }
904 // A barrier, which is one instruction or none depending on the ordering. Written
905 // by name because there is nothing about it a rule could be proved against, the
906 // way there is nothing to prove about the address of a symbol.
907 Opcode::Fence => {
908 self.barrier(inst)?;
909 continue;
910 }
911 // A hint, written by name for the reason a barrier is and one step further: not
912 // only is there no equality for a proof to discharge, there is nothing about the
913 // program around it either. Which of the four instructions it is comes out of the
914 // number the builtin was given, which is beside the instruction rather than in it.
915 Opcode::Prefetch => {
916 self.hint(inst)?;
917 continue;
918 }
919 // Stopping, written by name for the first half of the barrier's reason: it
920 // computes nothing, so there is no term for a rule to replace, and what makes it
921 // right is what the operating system does with the fault rather than anything a
922 // proof over bitvectors could discharge.
923 Opcode::Trap => {
924 self.trap(inst);
925 continue;
926 }
927 // A compare and exchange, which is written by name because it produces two values
928 // and a rule produces one. The replacement of a rule is one term, a term names the
929 // value an instruction computes, and there is no way in that language to say that
930 // an instruction leaves an answer in one place and a yes or no in another.
931 Opcode::Cmpxchg => {
932 self.exchange(inst)?;
933 continue;
934 }
935 // A read modify write, which is written by name for a different reason: it produces
936 // one value, so a rule could name it, and what it does is not in the head a rule
937 // matches on. Every one of the thirteen operations is the same opcode at the same
938 // type and differs only in what is carried beside it, so one pattern would be all
939 // thirteen patterns. Of the thirteen only the three with an instruction reach here,
940 // since `crate::retry` turned the rest into loops a long way above this.
941 Opcode::AtomicRmw => {
942 self.modify(inst)?;
943 continue;
944 }
945 // An `asm` statement, whose lowering is its template and there is no term for a
946 // string. Written by name for the reason a barrier is, and before the x87 arm
947 // below so that an `asm` holding a `long double` is refused as the `asm` it is
948 // rather than as an instruction nothing computes.
949 Opcode::InlineAsm => {
950 self.assembly(inst)?;
951 continue;
952 }
953 // Anything at all with an eighty bit float in it, which is the one arm here
954 // chosen by a type rather than by an opcode, because what makes these different
955 // is not what they do but where the value is. A `long double` has no register,
956 // so it has no name in `crate::term` and no rule could bind one: every one of
957 // these is a group of instructions over a frame slot, written out below.
958 //
959 // Last of the arms, so that a call and a return with one of these in them reach
960 // the convention first and are refused by it, which is the truer answer: what is
961 // wrong there is where the value has to travel and not that nothing can compute
962 // it.
963 _ if self.touches_x87(inst) => {
964 self.x87(inst)?;
965 continue;
966 }
967 _ => {}
968 }
969 let matched = matched.ok_or_else(|| self.unsupported(inst))?;
970 self.emit(inst, &matched)?;
971 // After it is built rather than when it matched, so that what is recorded is the rules
972 // this function was lowered by and not the rules something was tried with.
973 self.fired.mark(matched.rule);
974 }
975 self.edges(block, out)
976 }
977
978 /// One call, which is built from the convention rather than matched against the table for the
979 /// same reason the arguments of the function itself are.
980 ///
981 /// The arguments are read before the call is built, which is what materializes a constant
982 /// argument into a register, since no call passes an immediate.
983 ///
984 /// A call to a name and a call through an address are both here, and what tells them apart is
985 /// the opcode rather than whether a callee was recorded, which is the same thing the verifier
986 /// reads. Through an address the first operand is the address and the arguments are the ones
987 /// behind it, and everything after that is the same: where each argument goes, where the value
988 /// comes back and which registers are gone across it are the convention's answers and the
989 /// convention does not ask what is being called.
990 fn called(&mut self, inst: Inst) -> Result<(), Unsupported> {
991 let data = &self.source[inst];
992 let Extra::Call(info) = data.extra else { return Err(self.unsupported(inst)) };
993 let info = self.source[info];
994 let indirect = data.opcode == Opcode::CallIndirect;
995
996 let values: Vec<Value> = self.source[data.args].to_vec();
997 let callee = if indirect {
998 let &address = values.first().ok_or_else(|| self.unsupported(inst))?;
999 abi::Callee::Through(self.reg_of(address)?)
1000 } else {
1001 abi::Callee::Named(info.callee.ok_or_else(|| self.unsupported(inst))?)
1002 };
1003
1004 // What the ABI asks of each argument, read out before any of them is, because reading one
1005 // borrows the function this is a table in. The ones the signature names are the signature's
1006 // answer and the ones behind them are the call's, which is where a structure passed to a
1007 // variadic callee by value says that its bytes travel: there is no parameter to say it on.
1008 let signature = &self.source[info.signature];
1009 let variadic = signature.variadic;
1010 let named: Vec<Abi> = signature.params.iter().map(|param| param.abi).collect();
1011 let beyond: Vec<Abi> = self.source[info.varargs].to_vec();
1012 // Every value that comes back and not only the first. A structure small enough to travel
1013 // in registers comes back in up to two of them, and which register each half is in is the
1014 // convention's answer, which is why the whole list goes to the same place the arguments do
1015 // rather than to a rule.
1016 let returns: Vec<Type> = signature.return_types().collect();
1017
1018 let mut args = Vec::with_capacity(values.len());
1019 for (index, value) in values.into_iter().skip(usize::from(indirect)).enumerate() {
1020 let abi = named.get(index).or_else(|| beyond.get(index - named.len()));
1021 let abi = abi.copied().unwrap_or_default();
1022 let ty = self.source[value].ty;
1023 // What travels for an eighty bit value is its bytes, so what the call is handed is
1024 // where they are rather than a register they are in, and there is no register they
1025 // could be in. Everything else about it is a sixteen byte object passed by value and
1026 // is built by the same code.
1027 let reg =
1028 if abi::on_the_stack(ty) { self.x87_slot(value) } else { self.reg_of(value)? };
1029 args.push(abi::Passing { ty, reg, abi });
1030 }
1031 let block = self.at.expect("a block is being filled");
1032 let what = abi::Calling { callee, args: &args, returns: &returns, variadic };
1033 let made = abi::call(&mut self.out, block, &what, self.conv, self.names)
1034 .map_err(|refused| Unsupported::Call { inst, refused })?;
1035 let calls = &mut self.stack.calls;
1036 *calls = Some(calls.unwrap_or(0).max(made.outgoing));
1037 // An eighty bit value came back on the x87 stack, and the one thing that has to happen
1038 // before anything else touches that stack is taking it off. So the `fstp` goes here, in
1039 // front of everything the block does next, and after it the value is in its slot and is
1040 // read the way every other one is.
1041 let results: Vec<Value> = self.source[inst].results().collect();
1042 if let [result] = results[..] {
1043 if abi::on_the_stack(self.source[result].ty) {
1044 let span = self.source.span(inst);
1045 let into = self.x87_slot(result);
1046 let into = self.through(into);
1047 self.x87_at("fstp_t", span, into);
1048 return Ok(());
1049 }
1050 }
1051 for (result, ®) in results.into_iter().zip(&made.results) {
1052 self.regs[result.index()] = Some(reg);
1053 }
1054 Ok(())
1055 }
1056
1057 /// The pointer a function returning through memory was handed, or nothing in a function that
1058 /// was not.
1059 ///
1060 /// It is the first parameter and the signature is what says so, since in the IR it is an
1061 /// ordinary pointer and reads like one everywhere in the body. A function with a signature
1062 /// like that and no entry block has nothing to give back and no body to give it back from.
1063 fn sret(&self) -> Option<Value> {
1064 let first = self.source.signature().params.first()?;
1065 if !matches!(first.abi, Abi::Sret { .. }) {
1066 return None;
1067 }
1068 self.source[self.source.entry()?].params.first().copied()
1069 }
1070
1071 /// One `return` the convention has to write, as the place each value has to be in by the end.
1072 ///
1073 /// One pseudo per value, each a read constrained to a return register, which is what a return
1074 /// of one value already is and is the whole of what either does. The `ret` itself comes from
1075 /// the epilogue for both, long after this, because the frame has to be given back first.
1076 ///
1077 /// The two register files are counted separately, so a structure of a `double` and a `long`
1078 /// leaves the `double` in the first vector register and the `long` in the first integer one
1079 /// rather than in the second of either. That is the same walk `rucc_codegen::abi` makes on
1080 /// the other side of the call, which is what makes the two ends agree.
1081 ///
1082 /// A function whose answer went through memory gives back the address it was handed, in front
1083 /// of nothing else, because a signature that returns that way returns nothing else. That the
1084 /// caller already knows the address is not enough: it is allowed to read the register instead,
1085 /// and a caller that does gets whatever the allocator last left there. In a leaf function that
1086 /// is usually the right answer by accident, and one call in the body is enough to make it a
1087 /// wild pointer, which is why this is written rather than left to luck.
1088 ///
1089 /// Where everything goes is worked out before anything is written, so a return this cannot
1090 /// make leaves no half of one behind.
1091 /// Whether what a `return` gives back is the one value that goes back on the x87 stack.
1092 fn gives_back_x87(&self, inst: Inst) -> bool {
1093 let [value] = self.source[self.source[inst].args] else { return false };
1094 abi::on_the_stack(self.source[value].ty)
1095 }
1096
1097 fn returned(&mut self, inst: Inst) -> Result<(), Unsupported> {
1098 let values: Vec<Value> = self.source[self.source[inst].args].to_vec();
1099 let (mut ints, mut floats) = (0usize, 0usize);
1100 let mut parts = Vec::with_capacity(values.len() + 1);
1101 // An eighty bit value goes back on the x87 stack, which is where the convention says it is
1102 // and is the one place a value is left rather than put in a register. So the whole of the
1103 // return is an `fld` of its slot, and the stack it leaves the value on is not empty at the
1104 // `ret`, which is the one time in this file that is true and is what the convention asks
1105 // for. What comes after is the epilogue, which gives the frame back and touches nothing in
1106 // the unit.
1107 if let [value] = values[..] {
1108 let ty = self.source[value].ty;
1109 if abi::on_the_stack(ty) && self.sret().is_none() {
1110 let span = self.source.span(inst);
1111 let from = self.x87_slot(value);
1112 let from = self.through(from);
1113 self.x87_at("fld_t", span, from);
1114 return Ok(());
1115 }
1116 }
1117 for value in self.sret().into_iter().chain(values) {
1118 let ty = self.source[value].ty;
1119 let at = if crate::term::in_vector_file(ty) { &mut floats } else { &mut ints };
1120 // Why it cannot come back, and not only that it cannot. A type that travels nowhere
1121 // says so itself, and a type that travels perfectly well ran out of registers.
1122 let missing = abi::refuses(ty).unwrap_or(Missing::NoRoom);
1123 let name = abi::ret_of(ty, *at).ok_or(Unsupported::Returned { inst, missing })?;
1124 *at += 1;
1125 // The register is the target's answer and not one worked out here, the same as it is
1126 // for a return of one value, so that both halves of a pair and every rule that writes
1127 // half of one are reading the same table.
1128 let opcode = name.strip_prefix(PREFIX).expect("a machine instruction of this target");
1129 let form = x86_64::form(opcode).ok_or_else(|| self.unsupported(inst))?;
1130 let [desc] = form.operands() else { return Err(self.unsupported(inst)) };
1131 parts.push((self.names.intern(name), self.reg_of(value)?, *desc));
1132 }
1133
1134 let block = self.at.expect("a block is being filled");
1135 let span = self.source.span(inst);
1136 for (opcode, reg, desc) in parts {
1137 let operand = mir::Operand {
1138 reg,
1139 class: desc.class,
1140 role: desc.role,
1141 constraint: desc.constraint,
1142 };
1143 self.out.build(block, mir::Opcode::new(opcode)).at(span).operand(operand).finish();
1144 }
1145 Ok(())
1146 }
1147
1148 /// One `alloca`: the bytes it asks for go on the list the frame is laid out from, and the
1149 /// address of them is one instruction.
1150 ///
1151 /// The instruction is a `lea` off the stack pointer, which is the one register that reaches
1152 /// the frame in every function, and its displacement is left at nothing because there is no
1153 /// frame yet. Which instruction is waiting for which local is remembered, and
1154 /// [`crate::finish`] fills the numbers in after [`crate::frame::Frame`] has placed them.
1155 ///
1156 /// There is deliberately no rule for `alloca` and no name for one in [`crate::term`], and
1157 /// that is what stops it being folded into something else. An operand shown as the
1158 /// instruction that computed it is offered to the matcher by its name, so an `alloca` with no
1159 /// name is one no pattern can reach past, and the address it computes is always in a register
1160 /// by the time anything reads it.
1161 fn reserve(&mut self, inst: Inst) -> Result<(), Unsupported> {
1162 let data = &self.source[inst];
1163 // A variable length array carries the size it wants as an operand rather than in the
1164 // instruction, which is the whole of what tells the two apart here.
1165 if let Some(&size) = self.source[data.args].first() {
1166 return self.grow(inst, size);
1167 }
1168 let Extra::Mem(mem) = data.extra else { return Err(self.unsupported(inst)) };
1169 let info = self.source[mem];
1170 let size = u32::try_from(info.size)
1171 .map_err(|_| Unsupported::Dynamic { inst, growing: Growing::Huge })?;
1172 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
1173
1174 // At least one, because the frame divides by the alignment and an object with no
1175 // alignment at all is one the front end had nothing to say about rather than one that may
1176 // go anywhere.
1177 let index = self.stack.locals.len();
1178 self.stack.locals.push(Local { size, align: info.align.max(1) });
1179
1180 let block = self.at.expect("a block is being filled");
1181 let reg = self.new_reg(result);
1182 let span = self.source.span(inst);
1183 let lea = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.lea)));
1184 let sp = mir::Operand::read(mir::Reg::physical(self.conv.stack_pointer), self.gpr);
1185 let made =
1186 self.out.build(block, lea).at(span).def(reg, self.gpr).mem(mir::Mem::at(sp)).finish();
1187 self.stack.addresses.push((made, index));
1188 Ok(())
1189 }
1190
1191 /// The other kind of `alloca`: one whose size the function does not know until it runs, which
1192 /// is what a variable length array is.
1193 ///
1194 /// Nothing about it is a slot the frame laid out, because the frame is laid out once and this
1195 /// happens as often as control reaches the declaration. The bytes come off the stack pointer
1196 /// where the declaration stands, which is two instructions:
1197 ///
1198 /// ```text
1199 /// sub sp, bytes the stack pointer moves down over the memory, which is what takes it
1200 /// lea reg, [sp+n] where the memory starts, which is above the outgoing argument area
1201 /// ```
1202 ///
1203 /// The displacement is left at nothing for the reason the constant kind leaves its own at
1204 /// nothing, and for a different number: that area belongs to the arguments of whatever this
1205 /// function calls, it stays at the bottom of the frame wherever the bottom has moved to, and
1206 /// how big it is is not known until every call in the function has been seen.
1207 ///
1208 /// The bytes are already a multiple of the stack pointer's alignment by the time they arrive,
1209 /// because [`crate::expand::rounds`] rounded them up in the IR, so nothing here has to mask the
1210 /// stack pointer afterwards and the stack pointer stays somewhere a call can be made from.
1211 ///
1212 /// Two instructions here and not always two in the finished function. On a command line that
1213 /// asked for the stack to be touched a page at a time, the subtraction becomes a loop that
1214 /// walks the same distance a page at a time, which [`crate::finish`] writes. That is why the
1215 /// instruction is written down in [`Stack::grown`] as well as left where it is.
1216 ///
1217 /// Refused for an array wanting more alignment than the convention leaves the stack pointer
1218 /// with. Forcing that would be a second rounding of a register the frame already rounded, and
1219 /// after it no constant reaches the rest of the frame from anywhere. See `Growing` in
1220 /// [`crate::frame`].
1221 fn grow(&mut self, inst: Inst, size: Value) -> Result<(), Unsupported> {
1222 let data = &self.source[inst];
1223 let Extra::Mem(mem) = data.extra else { return Err(self.unsupported(inst)) };
1224 let info = self.source[mem];
1225 if info.align > self.conv.stack_align {
1226 return Err(Unsupported::Dynamic { inst, growing: Growing::Aligned });
1227 }
1228 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
1229 let bytes = self.reg_of(size)?;
1230
1231 let block = self.at.expect("a block is being filled");
1232 let span = self.source.span(inst);
1233 let stack = mir::Reg::physical(self.conv.stack_pointer);
1234 let grow = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.grow)));
1235 let took = self
1236 .out
1237 .build(block, grow)
1238 .at(span)
1239 .operand(mir::Operand::write(stack, self.gpr))
1240 .operand(mir::Operand::read(stack, self.gpr))
1241 .operand(mir::Operand::read(bytes, self.gpr))
1242 .finish();
1243 self.stack.grown.push(took);
1244
1245 let reg = self.new_reg(result);
1246 let lea = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.lea)));
1247 let sp = mir::Operand::read(stack, self.gpr);
1248 let made =
1249 self.out.build(block, lea).at(span).def(reg, self.gpr).mem(mir::Mem::at(sp)).finish();
1250 self.stack.dynamic.push(made);
1251 self.stack.grown_at.get_or_insert(inst);
1252 Ok(())
1253 }
1254
1255 /// Where the stack pointer is, kept so that something later can put it back.
1256 ///
1257 /// One move out of the stack pointer and one move into it, which is the whole of what the two
1258 /// halves are. What makes them worth writing is where the front end puts them: a scope holding
1259 /// a variable length array saves the stack pointer as it opens and puts it back as it closes,
1260 /// so a loop declaring one takes its bytes once round rather than once per iteration, and a
1261 /// jump out of the scope gives the bytes back on the way out.
1262 ///
1263 /// The value travels in an ordinary register the allocator hands out, so it may be spilled like
1264 /// any other, and a spill slot in a frame that grows is reached through the frame pointer,
1265 /// which is exactly the register that still means something after the stack pointer has moved.
1266 fn stack_pointer(&mut self, inst: Inst, into: bool) -> Result<(), Unsupported> {
1267 let data = &self.source[inst];
1268 let block = self.at.expect("a block is being filled");
1269 let span = self.source.span(inst);
1270 let stack = mir::Reg::physical(self.conv.stack_pointer);
1271 let mov = x86_64::FRAME.moves(self.gpr).expect("a class the target says how to move").mov;
1272 let mov = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{mov}")));
1273 let (write, read) = if into {
1274 let &saved = self.source[data.args].first().ok_or_else(|| self.unsupported(inst))?;
1275 (stack, self.reg_of(saved)?)
1276 } else {
1277 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
1278 (self.new_reg(result), stack)
1279 };
1280 self.out
1281 .build(block, mov)
1282 .at(span)
1283 .operand(mir::Operand::write(write, self.gpr))
1284 .operand(mir::Operand::read(read, self.gpr))
1285 .finish();
1286 // Only the write is a move of the stack pointer, and it is the one that makes the frame a
1287 // growing one. A read of it in a function that never writes it back is a function that
1288 // asked where the stack was and did nothing with the answer.
1289 if into {
1290 self.stack.grown_at.get_or_insert(inst);
1291 }
1292 Ok(())
1293 }
1294
1295 /// Whether an instruction has an eighty bit float anywhere in it.
1296 ///
1297 /// Producing one and reading one are the same question here, because what makes one of these
1298 /// different from every other instruction is not the operation but where the value is. A
1299 /// `long double` is on the x87 stack while it is being worked on and in a frame slot the rest
1300 /// of the time, and neither of those is somewhere the operand of a rule could point.
1301 fn touches_x87(&self, inst: Inst) -> bool {
1302 let data = &self.source[inst];
1303 data.results().any(|value| on_x87(self.source[value].ty))
1304 || self.source[data.args].iter().any(|&arg| on_x87(self.source[arg].ty))
1305 }
1306
1307 /// Everything that happens to an eighty bit float, as the group of instructions it is.
1308 ///
1309 /// The first six move one, and every one of those is a load, a store, or a load and a store at
1310 /// two different formats, because that is the whole of what this machine converts with: the
1311 /// x87 has no instruction that turns one thing on its stack into another, so a widening is
1312 /// `fld` of the narrow format and a narrowing is `fstp` of it.
1313 ///
1314 /// The rest work on one, and they are here rather than in a rule for the same reason the six
1315 /// are. An add is a push, a push, the add and a pop, and what passes between those four is the
1316 /// top of a stack nothing allocates from, so there is no value in the middle of the group for
1317 /// a pattern to bind or a replacement to name. The comparison is the same shape with its last
1318 /// two instructions folded into one opcode, which is where the byte it produces comes from.
1319 ///
1320 /// Every group leaves the stack as empty as it found it, which is what `spec/10-backend.md`
1321 /// section 10.8 asks of one and is why nothing in this file has to track a depth: each push
1322 /// below is answered by a pop a line or two later, so no two groups can ever be looking at
1323 /// the same eight registers.
1324 fn x87(&mut self, inst: Inst) -> Result<(), Unsupported> {
1325 match self.source[inst].opcode {
1326 Opcode::Load => self.x87_load(inst),
1327 Opcode::Store => self.x87_store(inst),
1328 Opcode::FPExt => self.x87_widen(inst),
1329 Opcode::FPTrunc => self.x87_narrow(inst),
1330 Opcode::SIToFP => self.x87_from_signed(inst),
1331 Opcode::FPToSI => self.x87_to_signed(inst),
1332 Opcode::FAdd => self.x87_arith(inst, "fadd_p"),
1333 Opcode::FSub => self.x87_arith(inst, "fsubr_p"),
1334 Opcode::FMul => self.x87_arith(inst, "fmul_p"),
1335 Opcode::FDiv => self.x87_arith(inst, "fdivr_p"),
1336 Opcode::FNeg => self.x87_flip(inst),
1337 Opcode::FCmp => self.x87_compare(inst),
1338 Opcode::FConst => self.x87_const(inst),
1339 _ => Err(self.unsupported(inst)),
1340 }
1341 }
1342
1343 /// The eighty bit parameters of a block, copied out of the addresses an edge handed over and
1344 /// into slots of the block's own.
1345 ///
1346 /// What crosses an edge for a value of this type is an address, because the value is sixteen
1347 /// bytes of the frame and no register holds any of it. The block cannot keep that address: a
1348 /// second edge into the same block hands over a second one, and a read after the block would
1349 /// then be a read of whichever edge was taken rather than of one place. So the block has a
1350 /// slot per parameter and the bytes are copied into it here, which is the move on an edge that
1351 /// every other type gets from the allocator.
1352 ///
1353 /// Every load runs before every store and the stores run backwards, so all of the values are
1354 /// on the x87 stack at once and nothing reads a slot another one has already written. That
1355 /// costs nothing in the ordinary case of one parameter and is what makes the back edge of a
1356 /// loop that swaps two of these work. It is also the reason for the limit: the stack is eight
1357 /// deep, and a block with more of these than that is refused rather than copied in an order
1358 /// that could be wrong.
1359 fn settle(&mut self, block: Block, arriving: &[(Value, mir::Reg)]) -> Result<(), Unsupported> {
1360 let Some(&(first, _)) = arriving.first() else { return Ok(()) };
1361 if arriving.len() > X87_DEPTH {
1362 let ty = self.source[first].ty;
1363 return Err(Unsupported::Phi { block, count: arriving.len(), ty });
1364 }
1365 // A block parameter comes from no instruction, so what this points at is the first thing
1366 // in the block, which is where a reader looking for the copy would look.
1367 let first_inst = self.source.insts(block).next();
1368 let span = first_inst.map_or(Span::DUMMY, |it| self.source.span(it));
1369 for &(_, reg) in arriving {
1370 let from = self.through(reg);
1371 self.x87_at("fld_t", span, from);
1372 }
1373 for &(param, _) in arriving.iter().rev() {
1374 let into = self.x87_slot(param);
1375 let into = self.through(into);
1376 self.x87_at("fstp_t", span, into);
1377 }
1378 Ok(())
1379 }
1380
1381 /// The frame slot an eighty bit value lives in, as its address in a fresh register.
1382 ///
1383 /// The slot is the value's for the whole function and is taken the first time somebody asks.
1384 /// The address is worked out again every time, which is a `lea` per use and is deliberate: one
1385 /// address kept in a register from the definition to the last use would hold a general purpose
1386 /// register open across everything in between, and a function with a handful of these in it
1387 /// would spend its registers on addresses of things rather than on things.
1388 fn x87_slot(&mut self, value: Value) -> mir::Reg {
1389 // An argument of the function has a slot already and it is the caller's. The convention
1390 // puts the bytes in the argument area and hands over where they are, so the address that
1391 // arrived is the answer and no second copy of the value is made. Nothing ever writes to a
1392 // value of this type once it exists, so nothing writes to the caller's copy either. A
1393 // parameter of any other block is not this: what arrived there is an address a predecessor
1394 // chose, [`Lowering::settle`] has already copied the bytes out of it, and the slot those
1395 // bytes landed in is the one below.
1396 let entry = self.source.entry();
1397 if let (Def::Param { block, .. }, Some(reg)) =
1398 (self.source[value].def, self.regs[value.index()])
1399 {
1400 if entry == Some(block) {
1401 return reg;
1402 }
1403 }
1404 let index = match self.slots[value.index()] {
1405 Some(index) => index,
1406 None => {
1407 let index = self.stack.locals.len();
1408 self.stack.locals.push(Local { size: X87_BYTES, align: X87_BYTES });
1409 self.slots[value.index()] = Some(index);
1410 index
1411 }
1412 };
1413 let block = self.at.expect("a block is being filled");
1414 self.frame_address(block, index)
1415 }
1416
1417 /// The bytes a value crosses between a register and the x87 stack through, as their address
1418 /// in a fresh register.
1419 fn x87_crossing(&mut self) -> mir::Reg {
1420 let index = match self.crossing {
1421 Some(index) => index,
1422 None => {
1423 let index = self.stack.locals.len();
1424 self.stack.locals.push(Local { size: X87_CROSSING, align: X87_CROSSING });
1425 self.crossing = Some(index);
1426 index
1427 }
1428 };
1429 let block = self.at.expect("a block is being filled");
1430 self.frame_address(block, index)
1431 }
1432
1433 /// The two control words, as the address of the first of them in a fresh register.
1434 fn x87_control(&mut self) -> mir::Reg {
1435 let index = match self.control {
1436 Some(index) => index,
1437 None => {
1438 let index = self.stack.locals.len();
1439 self.stack.locals.push(Local { size: 4, align: 4 });
1440 self.control = Some(index);
1441 index
1442 }
1443 };
1444 let block = self.at.expect("a block is being filled");
1445 self.frame_address(block, index)
1446 }
1447
1448 /// An address held in a register, as the addressing mode that reaches it.
1449 fn through(&self, reg: mir::Reg) -> mir::Mem {
1450 mir::Mem::at(mir::Operand::read(reg, self.gpr))
1451 }
1452
1453 /// One instruction of a group, which names an address and nothing else.
1454 ///
1455 /// Every x87 instruction that moves a value is one of these. What it does to the stack is in
1456 /// the mnemonic rather than in an operand, so there is no register to write down and no
1457 /// register the allocator gets a say in.
1458 fn x87_at(&mut self, name: &str, span: Span, at: mir::Mem) {
1459 let block = self.at.expect("a block is being filled");
1460 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
1461 self.out.build(block, opcode).at(span).mem(at).finish();
1462 }
1463
1464 /// One instruction of a group that names nothing at all.
1465 ///
1466 /// The arithmetic is these. Both of an add's operands are already on the stack when it runs
1467 /// and so is where the answer goes, and the stack is not somewhere an instruction says, so
1468 /// `faddp` has an argument in the assembler's syntax and nothing here for the argument to come
1469 /// from. What it works on is which two pushes came before it, which is a fact about the order
1470 /// of the group and is why the group is written in one place.
1471 fn x87_only(&mut self, name: &str, span: Span) {
1472 let block = self.at.expect("a block is being filled");
1473 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
1474 self.out.build(block, opcode).at(span).finish();
1475 }
1476
1477 /// A `load` of a `long double`: onto the stack from where it was, and off it into the slot.
1478 ///
1479 /// Two instructions rather than the two general purpose moves the same sixteen bytes would
1480 /// take, because `fld` and `fstp` at this format neither convert nor look: the value goes on
1481 /// in the format it was already in and comes back off in it, so a signalling NaN stays one
1482 /// and nothing is raised. Which is what makes this a copy at all.
1483 fn x87_load(&mut self, inst: Inst) -> Result<(), Unsupported> {
1484 let (args, result) = self.ends(inst)?;
1485 let &address = args.first().ok_or_else(|| self.unsupported(inst))?;
1486 let span = self.source.span(inst);
1487 let from = self.reg_of(address)?;
1488 let from = self.through(from);
1489 let into = self.x87_slot(result);
1490 let into = self.through(into);
1491 self.x87_at("fld_t", span, from);
1492 self.x87_at("fstp_t", span, into);
1493 Ok(())
1494 }
1495
1496 /// A `store` of a `long double`: the same pair the other way round.
1497 fn x87_store(&mut self, inst: Inst) -> Result<(), Unsupported> {
1498 let args = self.source[self.source[inst].args].to_vec();
1499 let [value, address] = args[..] else { return Err(self.unsupported(inst)) };
1500 let span = self.source.span(inst);
1501 let from = self.x87_slot(value);
1502 let from = self.through(from);
1503 let into = self.reg_of(address)?;
1504 let into = self.through(into);
1505 self.x87_at("fld_t", span, from);
1506 self.x87_at("fstp_t", span, into);
1507 Ok(())
1508 }
1509
1510 /// A `float`, a `double` or an integer becoming a `long double`.
1511 ///
1512 /// Through memory, because the x87 reads memory and nothing else: the value is in a register
1513 /// the machine has and the unit has no way to be handed one, so it is written to the crossing
1514 /// bytes and loaded back at the format that widens it. Every one of these is exact. Sixty four
1515 /// bits of significand and fifteen of exponent hold every `float`, every `double` and every
1516 /// sixty four bit integer outright, so none of the four can round and none can raise.
1517 fn x87_across(
1518 &mut self,
1519 inst: Inst,
1520 put: &'static str,
1521 class: RegClass,
1522 get: &'static str,
1523 ) -> Result<(), Unsupported> {
1524 let (args, result) = self.ends(inst)?;
1525 let &source = args.first().ok_or_else(|| self.unsupported(inst))?;
1526 let span = self.source.span(inst);
1527 let value = self.reg_of(source)?;
1528 let across = self.x87_crossing();
1529 let across = self.through(across);
1530 let into = self.x87_slot(result);
1531 let into = self.through(into);
1532
1533 let block = self.at.expect("a block is being filled");
1534 let store = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{put}")));
1535 self.out.build(block, store).at(span).uses(value, class).mem(across).finish();
1536 self.x87_at(get, span, across);
1537 self.x87_at("fstp_t", span, into);
1538 Ok(())
1539 }
1540
1541 /// A `long double` becoming a `float`, a `double` or an integer.
1542 ///
1543 /// Through memory for the reason above and in the same three instructions backwards. The two
1544 /// that go to a float round to nearest, which is what the control word says unless somebody
1545 /// has changed it and is what C wants. The two that go to an integer do not, which is why they
1546 /// do not come here.
1547 fn x87_back(
1548 &mut self,
1549 inst: Inst,
1550 put: &'static str,
1551 get: &'static str,
1552 class: RegClass,
1553 ) -> Result<(), Unsupported> {
1554 let (args, result) = self.ends(inst)?;
1555 let &source = args.first().ok_or_else(|| self.unsupported(inst))?;
1556 let span = self.source.span(inst);
1557 let from = self.x87_slot(source);
1558 let from = self.through(from);
1559 let across = self.x87_crossing();
1560 let across = self.through(across);
1561
1562 self.x87_at("fld_t", span, from);
1563 self.x87_at(put, span, across);
1564 let block = self.at.expect("a block is being filled");
1565 let reg = self.new_reg(result);
1566 let load = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{get}")));
1567 self.out.build(block, load).at(span).def(reg, class).mem(across).finish();
1568 Ok(())
1569 }
1570
1571 /// An `fpext` up to a `long double`, which is the only direction this machine has one in.
1572 fn x87_widen(&mut self, inst: Inst) -> Result<(), Unsupported> {
1573 let sse = self.conv.sse_class;
1574 match self.source[self.narrow(inst)?].ty.bits() {
1575 32 => self.x87_across(inst, "movss_mr", sse, "fld_s"),
1576 64 => self.x87_across(inst, "movsd_mr", sse, "fld_l"),
1577 _ => Err(self.unsupported(inst)),
1578 }
1579 }
1580
1581 /// An `fptrunc` down from a `long double`, which is the other direction of the same.
1582 fn x87_narrow(&mut self, inst: Inst) -> Result<(), Unsupported> {
1583 let sse = self.conv.sse_class;
1584 let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
1585 match self.source[result].ty.bits() {
1586 32 => self.x87_back(inst, "fstp_s", "movss_rm", sse),
1587 64 => self.x87_back(inst, "fstp_l", "movsd_rm", sse),
1588 _ => Err(self.unsupported(inst)),
1589 }
1590 }
1591
1592 /// A `sitofp` up to a `long double`.
1593 ///
1594 /// Thirty two bits and sixty four, and nothing narrower, because C widens an integer to `int`
1595 /// before it converts one and the front end writes that widening down. An unsigned integer is
1596 /// not here at all: `fild` reads its operand as signed, so a value above the signed range
1597 /// comes back short by two to the sixty fourth and has to be added back, which is arithmetic
1598 /// rather than a move and waits with the rest of it.
1599 fn x87_from_signed(&mut self, inst: Inst) -> Result<(), Unsupported> {
1600 let gpr = self.gpr;
1601 match self.source[self.narrow(inst)?].ty.bits() {
1602 32 => self.x87_across(inst, "mov_mr_32", gpr, "fild_l"),
1603 64 => self.x87_across(inst, "mov_mr_64", gpr, "fild_ll"),
1604 _ => Err(self.unsupported(inst)),
1605 }
1606 }
1607
1608 /// An `fptosi` down from a `long double`, which is the one conversion here with no single
1609 /// instruction behind it.
1610 ///
1611 /// C cuts towards zero and the unit rounds the way its control word says, so the store that
1612 /// takes the value off the stack is wrapped in the control word being saved, changed and put
1613 /// back. Five instructions around the one that does the work, and three more moving the word
1614 /// through a register, because this machine has no way to OR a constant into memory at this
1615 /// width. The unit has a shorter answer in `fisttp`, and `spec/10-backend.md` section 10.8
1616 /// says why it is not used: it is SSE3, the x86-64 baseline is not, and there is nothing here
1617 /// that can gate an instruction on a feature yet.
1618 fn x87_to_signed(&mut self, inst: Inst) -> Result<(), Unsupported> {
1619 let (args, result) = self.ends(inst)?;
1620 let &source = args.first().ok_or_else(|| self.unsupported(inst))?;
1621 let (put, get) = match self.source[result].ty.bits() {
1622 32 => ("fistp_l", "mov_rm_32"),
1623 64 => ("fistp_ll", "mov_rm_64"),
1624 _ => return Err(self.unsupported(inst)),
1625 };
1626 let span = self.source.span(inst);
1627 let gpr = self.gpr;
1628 let from = self.x87_slot(source);
1629 let from = self.through(from);
1630 let across = self.x87_crossing();
1631 let across = self.through(across);
1632 let control = self.x87_control();
1633 let saved = self.through(control).plus(0);
1634 let cut = self.through(control).plus(2);
1635
1636 // The word the unit has now, into the first of the two slots and into a register, with the
1637 // rounding field turned to truncate on the way to the second.
1638 self.x87_at("fnstcw", span, saved);
1639 let block = self.at.expect("a block is being filled");
1640 let was = self.out.new_vreg(gpr);
1641 let read = mir::Opcode::new(self.names.intern("x64.mov_rm_16"));
1642 self.out.build(block, read).at(span).def(was, gpr).mem(saved).finish();
1643 let now = self.out.new_vreg(gpr);
1644 let set = mir::Opcode::new(self.names.intern("x64.or_ri_16"));
1645 // Two address, which is written out here rather than taken from the two shorthands
1646 // because the shorthands leave an operand unconstrained: this machine ORs into the
1647 // register it read, so the two have to be the same one and only the constraint says so.
1648 self.out
1649 .build(block, set)
1650 .at(span)
1651 .operand(mir::Operand::write(now, gpr).with(Constraint::Reuse(1)))
1652 .operand(mir::Operand::read(was, gpr))
1653 .imm(X87_TRUNCATE)
1654 .finish();
1655 let write = mir::Opcode::new(self.names.intern("x64.mov_mr_16"));
1656 self.out.build(block, write).at(span).uses(now, gpr).mem(cut).finish();
1657
1658 // The conversion itself, under the changed word, and then the word the unit had put back
1659 // before anything else runs.
1660 self.x87_at("fldcw", span, cut);
1661 self.x87_at("fld_t", span, from);
1662 self.x87_at(put, span, across);
1663 self.x87_at("fldcw", span, saved);
1664
1665 let block = self.at.expect("a block is being filled");
1666 let reg = self.new_reg(result);
1667 let load = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{get}")));
1668 self.out.build(block, load).at(span).def(reg, gpr).mem(across).finish();
1669 Ok(())
1670 }
1671
1672 /// A constant of this type, as the bits of it written into its slot.
1673 ///
1674 /// No x87 instruction at all, which is the surprise here. A slot holding an eighty bit value is
1675 /// the value, so a constant is ten bytes put where the value lives, and the unit never has to
1676 /// see it: whatever reads it will `fld` it out of the slot the way it reads any other one.
1677 ///
1678 /// Ten bytes in two goes, because the machine stores eight at a time and there is no store of
1679 /// an immediate to memory, so each half is put in a register first. The six bytes above the ten
1680 /// are left alone, since nothing reads them: they are the padding that makes the type sixteen
1681 /// wide and they are unspecified in the psABI rather than zero.
1682 ///
1683 /// The other way is a constant pool, an `fldt` of a symbol, and a relocation, which is what a
1684 /// compiler with somewhere to put a literal does. This back end has nowhere to put one yet, and
1685 /// four instructions in the frame is what that costs until it does.
1686 fn x87_const(&mut self, inst: Inst) -> Result<(), Unsupported> {
1687 let Extra::Imm(imm) = self.source[inst].extra else { return Err(self.unsupported(inst)) };
1688 let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
1689 let bits = self.source[imm].bits();
1690 let span = self.source.span(inst);
1691 let gpr = self.gpr;
1692 let slot = self.x87_slot(result);
1693 let low = self.through(slot).plus(0);
1694 let high = self.through(slot).plus(8);
1695
1696 let block = self.at.expect("a block is being filled");
1697 for (bytes, at, into) in
1698 [(bits as u64 as i64, low, "64"), (((bits >> 64) & 0xffff) as i64, high, "16")]
1699 {
1700 let held = self.out.new_vreg(gpr);
1701 let put = mir::Opcode::new(self.names.intern(&format!("{PREFIX}mov_ri_{into}")));
1702 self.out.build(block, put).at(span).def(held, gpr).imm(bytes).finish();
1703 let store = mir::Opcode::new(self.names.intern(&format!("{PREFIX}mov_mr_{into}")));
1704 self.out.build(block, store).at(span).uses(held, gpr).mem(at).finish();
1705 }
1706 Ok(())
1707 }
1708
1709 /// One arithmetic instruction on two eighty bit values, as the four it takes.
1710 ///
1711 /// The left operand is pushed first and the right one on top of it, so the left ends up
1712 /// underneath and the answer wanted is the one below against the top in that order. Which of
1713 /// the two mnemonics computes that is a question about the spelling rather than about the
1714 /// machine, and the two spellings disagree. Intel's `FSUBP ST(i), ST(0)` is `ST(i) - ST(0)`
1715 /// and is `DE E8+i`, and AT&T's `fsubp` is `DE E0+i`, which is the other subtraction. This
1716 /// compiler writes AT&T and encodes what gas encodes, so what it asks for here is `fsubr_p`
1717 /// and `fdivr_p`, and the `r` is not a reversal of anything the code generator decided.
1718 ///
1719 /// An addition and a multiplication have one form each and do not care, which is why a test
1720 /// that reads the mnemonic back would not have caught this and one that computes a subtraction
1721 /// and checks the answer does.
1722 ///
1723 /// The answer is left where the deeper of the two was and the shallower is gone, which is what
1724 /// the `p` on the mnemonic means, so one push has already been paid back by the time the
1725 /// `fstp` runs and the stack is level again after it.
1726 ///
1727 /// Nothing here is folded and nothing is reused. Two values that are the same value get two
1728 /// pushes of the same slot, and an operand that was just computed is read back out of the slot
1729 /// it was written to rather than left on the stack, which costs a store and a load per
1730 /// instruction in an expression. Keeping a partial result on the stack across the next
1731 /// instruction's operands means knowing how deep the stack is at every point in the block, and
1732 /// that is a different thing from writing a group.
1733 fn x87_arith(&mut self, inst: Inst, with: &'static str) -> Result<(), Unsupported> {
1734 let (args, result) = self.ends(inst)?;
1735 let [left, right] = args[..] else { return Err(self.unsupported(inst)) };
1736 let span = self.source.span(inst);
1737 let left = self.x87_slot(left);
1738 let left = self.through(left);
1739 let right = self.x87_slot(right);
1740 let right = self.through(right);
1741 let into = self.x87_slot(result);
1742 let into = self.through(into);
1743 self.x87_at("fld_t", span, left);
1744 self.x87_at("fld_t", span, right);
1745 self.x87_only(with, span);
1746 self.x87_at("fstp_t", span, into);
1747 Ok(())
1748 }
1749
1750 /// A negation, which is a push, the sign bit turned over and a pop.
1751 ///
1752 /// `fchs` does not read the value as a number, so this is right for a zero, for an infinity
1753 /// and for a NaN, and it raises nothing on any of them. Which is what C asks of a negation and
1754 /// is not what subtracting from zero would give: `0.0L - x` is a different answer at a
1755 /// negative zero and a signalling one at a NaN.
1756 fn x87_flip(&mut self, inst: Inst) -> Result<(), Unsupported> {
1757 let (args, result) = self.ends(inst)?;
1758 let &source = args.first().ok_or_else(|| self.unsupported(inst))?;
1759 let span = self.source.span(inst);
1760 let from = self.x87_slot(source);
1761 let from = self.through(from);
1762 let into = self.x87_slot(result);
1763 let into = self.through(into);
1764 self.x87_at("fld_t", span, from);
1765 self.x87_only("fchs", span);
1766 self.x87_at("fstp_t", span, into);
1767 Ok(())
1768 }
1769
1770 /// A comparison of two eighty bit values, as the two pushes and the one opcode that reads them.
1771 ///
1772 /// The right operand is pushed first and the left one on top of it, which is the other way
1773 /// round from the arithmetic and is because `fucomip` asks about the top against what is under
1774 /// it: the comparison this machine can do is the top's, so the value the predicate is about
1775 /// has to be the top. The pop that gets the loser off the stack and the byte that reads the
1776 /// flags are both inside the opcode, since what passes between those and the comparison is the
1777 /// flags and the flags are not something anything here can name.
1778 ///
1779 /// Which of the ten opcodes, and which way round, is the same table the vector comparisons
1780 /// match against in `rules/x86-64.rules`, and it has to stay the same table: a predicate that
1781 /// picked a different condition here than there would be a `long double` comparison that
1782 /// disagreed with the `double` comparison of the same two numbers, which is the one thing a
1783 /// wider format is not allowed to do.
1784 ///
1785 /// The always false and the always true are refused rather than folded into a constant,
1786 /// because a comparison this machine never has to do is one the optimizer should have removed
1787 /// and an instruction here that quietly agreed with it would hide that it did not.
1788 fn x87_compare(&mut self, inst: Inst) -> Result<(), Unsupported> {
1789 let Extra::FloatPred(pred) = self.source[inst].extra else {
1790 return Err(self.unsupported(inst));
1791 };
1792 let (args, result) = self.ends(inst)?;
1793 let [left, right] = args[..] else { return Err(self.unsupported(inst)) };
1794 // Two of the fourteen need a second byte and an instruction to put the two together,
1795 // because they are two conditions at once: an ordered equal is equal and not unordered,
1796 // and an unordered not equal is either. The opcode carries all of that and says here only
1797 // that it writes somewhere else as well.
1798 let (name, reversed, both) = match pred {
1799 FloatPred::Ogt => ("fucomip_set_a", false, false),
1800 FloatPred::Oge => ("fucomip_set_ae", false, false),
1801 FloatPred::Olt => ("fucomip_set_a", true, false),
1802 FloatPred::Ole => ("fucomip_set_ae", true, false),
1803 FloatPred::One => ("fucomip_set_ne", false, false),
1804 FloatPred::Ord => ("fucomip_set_np", false, false),
1805 FloatPred::Uno => ("fucomip_set_p", false, false),
1806 FloatPred::Ueq => ("fucomip_set_e", false, false),
1807 FloatPred::Ult => ("fucomip_set_b", false, false),
1808 FloatPred::Ule => ("fucomip_set_be", false, false),
1809 FloatPred::Ugt => ("fucomip_set_b", true, false),
1810 FloatPred::Uge => ("fucomip_set_be", true, false),
1811 FloatPred::Oeq => ("fucomip_set_e_and_np", false, true),
1812 FloatPred::Une => ("fucomip_set_ne_or_p", false, true),
1813 FloatPred::False | FloatPred::True => return Err(self.unsupported(inst)),
1814 };
1815 let (top, under) = if reversed { (right, left) } else { (left, right) };
1816
1817 let span = self.source.span(inst);
1818 let gpr = self.gpr;
1819 let under = self.x87_slot(under);
1820 let under = self.through(under);
1821 let top = self.x87_slot(top);
1822 let top = self.through(top);
1823 self.x87_at("fld_t", span, under);
1824 self.x87_at("fld_t", span, top);
1825
1826 let block = self.at.expect("a block is being filled");
1827 let reg = self.new_reg(result);
1828 // Taken before the instruction is started rather than inside it, since both come from the
1829 // same function being built and only one thing at a time may be adding to it.
1830 let spare = both.then(|| self.out.new_vreg(gpr));
1831 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
1832 let mut build = self.out.build(block, opcode).at(span).def(reg, gpr);
1833 if let Some(spare) = spare {
1834 build = build.def(spare, gpr);
1835 }
1836 build.finish();
1837 Ok(())
1838 }
1839
1840 /// The operands and the one result of an instruction that has exactly one.
1841 fn ends(&self, inst: Inst) -> Result<(&'a [Value], Value), Unsupported> {
1842 let data = &self.source[inst];
1843 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
1844 Ok((&self.source[data.args], result))
1845 }
1846
1847 /// The operand of a conversion, which is the end of it that is not the `long double`.
1848 fn narrow(&self, inst: Inst) -> Result<Value, Unsupported> {
1849 let args = &self.source[self.source[inst].args];
1850 args.first().copied().ok_or_else(|| self.unsupported(inst))
1851 }
1852
1853 /// One `va_start`, as the four fields of the list it was handed.
1854 ///
1855 /// Two of them are numbers this already knows, and each costs an instruction to put in a
1856 /// register before it can be stored, because the machine here has no store of an immediate to
1857 /// memory. The other two are addresses in the frame, and each is a `lea` [`crate::finish`]
1858 /// finishes: the save area is one of the function's own stack objects, and the caller's
1859 /// argument area is where the parameters that had no register came from, which is the same
1860 /// place and the same fixup a parameter past the sixth already uses.
1861 ///
1862 /// What is written is exactly the four fields [`crate::varargs`] describes, in the order they
1863 /// are laid out, so that reading this beside that table is the whole of the check.
1864 fn va_start(&mut self, inst: Inst) -> Result<(), Unsupported> {
1865 let Some(&list) = self.source[self.source[inst].args].first() else {
1866 return Err(self.unsupported(inst));
1867 };
1868 let started = self.varargs.ok_or_else(|| self.unsupported(inst))?;
1869 let list = self.reg_of(list)?;
1870 let block = self.at.expect("a block is being filled");
1871 let span = self.source.span(inst);
1872
1873 for (at, count) in
1874 [(varargs::GP_OFFSET, started.integers), (varargs::FP_OFFSET, started.floats)]
1875 {
1876 let held = self.out.new_vreg(self.gpr);
1877 let load = mir::Opcode::new(self.names.intern("x64.mov_ri_32"));
1878 self.out.build(block, load).at(span).def(held, self.gpr).imm(i64::from(count)).finish();
1879
1880 let store = mir::Opcode::new(self.names.intern("x64.mov_mr_32"));
1881 let mem = self.field(list, at);
1882 self.out.build(block, store).at(span).uses(held, self.gpr).mem(mem).finish();
1883 }
1884
1885 // The first argument the signature did not name, which is as far up the caller's argument
1886 // area as the ones it did name reached. Nothing here knows where that area is, so the
1887 // distance is recorded the way a parameter read out of it is and finished with it.
1888 let overflow = self.out.new_vreg(self.gpr);
1889 let lea = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.lea)));
1890 let sp = mir::Operand::read(mir::Reg::physical(self.conv.stack_pointer), self.gpr);
1891 let made = self
1892 .out
1893 .build(block, lea)
1894 .at(span)
1895 .def(overflow, self.gpr)
1896 .mem(mir::Mem::at(sp))
1897 .finish();
1898 self.stack.arguments.push((made, started.incoming));
1899
1900 let save = self.frame_address(block, started.save);
1901 for (at, held) in [(varargs::OVERFLOW, overflow), (varargs::SAVE_AREA, save)] {
1902 let store = mir::Opcode::new(self.names.intern("x64.mov_mr_64"));
1903 let mem = self.field(list, at);
1904 self.out.build(block, store).at(span).uses(held, self.gpr).mem(mem).finish();
1905 }
1906 Ok(())
1907 }
1908
1909 /// One field of a list, as the addressing mode that reaches it.
1910 fn field(&self, list: mir::Reg, at: i64) -> mir::Mem {
1911 let base = mir::Operand::read(list, self.gpr);
1912 mir::Mem::at(base).plus(i32::try_from(at).expect("a field of a list is a small offset"))
1913 }
1914
1915 /// The address of a name: one `lea` off the instruction pointer, with the name on it.
1916 ///
1917 /// The same instruction an `alloca` gets and for a related reason. An address that is not in
1918 /// the program is a `lea` of an addressing mode that names no register, and the mode carries
1919 /// the symbol so that [`rucc_asm`] can write it relative to `%rip` and leave the relocation
1920 /// for the assembler. Both halves of that already existed: the printer writes `sym(%rip)` and
1921 /// the encoder emits the relocation, because a call to a name the file does not define needed
1922 /// them first.
1923 ///
1924 /// One `mov` and not one `lea` when the name is one [`Elsewhere`] holds, because the distance
1925 /// the `lea` adds to the instruction pointer is a number only a link that puts the name in
1926 /// this program can work out, and the address of a function this file merely declares is not
1927 /// such a number. The load reads the address out of the slot the linker fills in instead. The
1928 /// linker turns it back into the `lea` when the name turns out to have been here all along,
1929 /// so this is not slower in the case that was already right.
1930 ///
1931 /// There is deliberately no name for this in [`crate::term`], which is what stops the address
1932 /// being folded into the instruction that reads it. Folding it is the right thing to do and
1933 /// is what turns a load of a global from two instructions into one, but it is a separate
1934 /// question about addressing modes and issue #282 is it. Until then the address is in a
1935 /// register before anything uses it, which is correct and one instruction longer.
1936 ///
1937 /// What this does not do is give the name anything to refer to. A module carries its globals
1938 /// and nothing writes them out, so a file that defines the variable it reads compiles to a
1939 /// reference the linker cannot resolve. Issue #293 is the other half.
1940 ///
1941 /// A thread-local variable is neither of the two above and is [`Self::thread_address`].
1942 fn address_of(&mut self, inst: Inst) -> Result<(), Unsupported> {
1943 let data = &self.source[inst];
1944 let Extra::Symbol(symbol) = data.extra else { return Err(self.unsupported(inst)) };
1945 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
1946 if self.elsewhere.thread(symbol) {
1947 return self.thread_address(inst, symbol, result);
1948 }
1949
1950 let block = self.at.expect("a block is being filled");
1951 let reg = self.new_reg(result);
1952 let span = self.source.span(inst);
1953 let (mnemonic, mem) = if self.elsewhere.holds(symbol) {
1954 (GOT_LOAD, mir::Mem::got(symbol))
1955 } else {
1956 (x86_64::FRAME.lea, mir::Mem::of(symbol))
1957 };
1958 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{mnemonic}")));
1959 self.out.build(block, opcode).at(span).def(reg, self.gpr).mem(mem).finish();
1960 Ok(())
1961 }
1962
1963 /// The address of a thread-local variable, which is this thread's copy of it.
1964 ///
1965 /// Neither instruction the ordinary case writes would mean anything here. There is no distance
1966 /// to the variable for a `lea` to add, because there is no variable: there is one copy of it per
1967 /// thread and they are at different addresses, so a link asked for the distance to the name
1968 /// refuses rather than picking one. And there is no address for a table slot to hold either, for
1969 /// the same reason.
1970 ///
1971 /// What is the same in every thread is where the variable sits inside the block of storage a
1972 /// thread gets, so that offset is what the link writes down, and the address of the running
1973 /// thread's block is what turns it into an address. x86-64 keeps that address in `%fs`, at the
1974 /// front of the block, so the whole of this is three instructions:
1975 ///
1976 /// ```text
1977 /// movq x@gottpoff(%rip), %off # how far into the block x sits, which the link fills in
1978 /// movq %fs:0, %tp # where this thread's block is, which only the machine knows
1979 /// addq %tp, %off # this thread's copy of x
1980 /// ```
1981 ///
1982 /// That is the initial exec model. It is one instruction longer than what gcc writes at `-O2`
1983 /// in an executable, which folds the addition into the instruction that uses the address, and
1984 /// the difference is issue #282 rather than anything about threads: nothing here folds an
1985 /// address into its reader yet. The link relaxes the first instruction into an immediate when it
1986 /// is making an executable, since it lays the blocks out and therefore knows the number, so the
1987 /// table slot costs nothing in the case that is common.
1988 ///
1989 /// It is not the most general model. A library loaded by `dlopen` gets its storage after the
1990 /// program is already running, and the block this reaches was laid out before it started, so
1991 /// the loader has to find room in that block for the library's variables. glibc keeps a little
1992 /// spare room for exactly this and a library that fits in it loads and runs; one that does not
1993 /// fails to load, with a message saying so. The model with no such limit calls `__tls_get_addr`
1994 /// and is what gcc writes under `-fPIC` by default, and it is issue #1104.
1995 ///
1996 /// So this is the model gcc writes under `-ftls-model=initial-exec`: right for an executable,
1997 /// right for a library the program is linked against, and a load that either works or is
1998 /// refused out loud for a library something opens later. What it is never is quietly wrong.
1999 fn thread_address(
2000 &mut self,
2001 inst: Inst,
2002 symbol: Symbol,
2003 result: Value,
2004 ) -> Result<(), Unsupported> {
2005 let block = self.at.expect("a block is being filled");
2006 let span = self.source.span(inst);
2007 let gpr = self.gpr;
2008 let load = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{GOT_LOAD}")));
2009
2010 let offset = self.out.new_vreg(gpr);
2011 self.out
2012 .build(block, load)
2013 .at(span)
2014 .def(offset, gpr)
2015 .mem(mir::Mem::thread(symbol))
2016 .finish();
2017 // The front of the block, which is the one thing on this machine that no instruction can
2018 // work out: `%fs` is not a register a program can read, and what it points at is a word
2019 // holding its own address, so reading through it at zero is how the address is come by.
2020 let pointer = self.out.new_vreg(gpr);
2021 let at = mir::Mem::in_segment(Segment::Fs, 0);
2022 self.out.build(block, load).at(span).def(pointer, gpr).mem(at).finish();
2023
2024 // Two address, spelled out for the reason `x87_to_int` gives: this machine adds into the
2025 // register it read, and only the constraint says the two are the same one.
2026 let reg = self.new_reg(result);
2027 let add = mir::Opcode::new(self.names.intern(&format!("{PREFIX}add_rr_64")));
2028 self.out
2029 .build(block, add)
2030 .at(span)
2031 .operand(mir::Operand::write(reg, gpr).with(Constraint::Reuse(1)))
2032 .operand(mir::Operand::read(offset, gpr))
2033 .operand(mir::Operand::read(pointer, gpr))
2034 .finish();
2035 Ok(())
2036 }
2037
2038 /// `&&label`, GNU's address of a label, which is the same `lea` a global gets against a place
2039 /// in this same function.
2040 ///
2041 /// What the two have in common is the whole of the instruction: an address worked out from
2042 /// where the instruction is, which is what `(%rip)` means and is the only way this compiler
2043 /// reaches anything. What they do not have in common is what fills the four bytes in. A
2044 /// global is a name, so the number is a relocation and the linker writes it. A block is a
2045 /// place in this function, so both ends are in one section and the number is known as soon as
2046 /// the blocks have been laid out, which is why `rucc_asm` fills it in the way it fills in a
2047 /// jump rather than leaving a relocation behind.
2048 ///
2049 /// Nothing here says the block is one control can arrive at. That is said by the
2050 /// [`Opcode::IndirectBr`] that reads the address, which lists every block it can arrive at,
2051 /// and by nothing else: an address on its own is a number.
2052 fn block_address(&mut self, inst: Inst) -> Result<(), Unsupported> {
2053 let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
2054 let Some(call) = self.source.successors(inst).next() else {
2055 return Err(self.unsupported(inst));
2056 };
2057 let block = self.at.expect("a block is being filled");
2058 let reg = self.new_reg(result);
2059 let span = self.source.span(inst);
2060 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.lea)));
2061 let mem = mir::Mem::block(self.out_block(call.block));
2062 self.out.build(block, opcode).at(span).def(reg, self.gpr).mem(mem).finish();
2063 Ok(())
2064 }
2065
2066 /// `goto *p`, GNU's computed goto, which is a jump through a register.
2067 ///
2068 /// Where it goes is not written here and cannot be. Every block it can arrive at is on the
2069 /// block this ends, the way every other arm is, and which of them the address holds is decided
2070 /// while the program runs. So this is one instruction with one operand, and the arms are
2071 /// copied across by [`Self::edges`] like anybody else's.
2072 fn indirect_branch(&mut self, inst: Inst) -> Result<(), Unsupported> {
2073 let data = &self.source[inst];
2074 let &address = self.source[data.args].first().ok_or_else(|| self.unsupported(inst))?;
2075 let reg = self.reg_of(address)?;
2076 let block = self.at.expect("a block is being filled");
2077 let span = self.source.span(inst);
2078 let name = x86_64::BRANCH.indirect;
2079 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
2080 self.out.build(block, opcode).at(span).operand(mir::Operand::read(reg, self.gpr)).finish();
2081 Ok(())
2082 }
2083
2084 /// `__builtin_frame_address` and `__builtin_return_address`, which are a walk up the chain of
2085 /// saved frame pointers and then one thing read at the end of it.
2086 ///
2087 /// Every frame that kept a frame pointer holds the caller's at the address the register points
2088 /// at, and the address that frame returns to one word above that, which is where the call
2089 /// instruction put it and where the prologue's push left it. So the walk is a load through the
2090 /// register for each link, the frame address is wherever the walk stopped, and the return
2091 /// address is one more load from a word above it. gcc 16.2.0 writes exactly this, measured on
2092 /// x86-64 at `-O2` for depths zero to three of both builtins.
2093 ///
2094 /// The function is given a frame pointer because of this, which is what [`Stack::walks_frames`]
2095 /// carries out to the layout. A depth of zero needs it as the answer and every depth above zero
2096 /// needs it as the start, so there is no case here where it is not wanted.
2097 ///
2098 /// How far the chain actually reaches is the program's business and not this one's. A caller
2099 /// compiled without a frame pointer has no link in it for the walk to follow, so a depth above
2100 /// zero is a promise about how the whole program was built. That is why gcc documents a nonzero
2101 /// depth as unsafe rather than as an answer, and why the depth is refused above a limit in
2102 /// `check/builtin/frame.rs` rather than walked as far as it says.
2103 fn frames(&mut self, inst: Inst) -> Result<(), Unsupported> {
2104 let data = &self.source[inst];
2105 let Extra::Depth(depth) = data.extra else { return Err(self.unsupported(inst)) };
2106 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
2107 let returning = data.opcode == Opcode::ReturnAddress;
2108 let block = self.at.expect("a block is being filled");
2109 let span = self.source.span(inst);
2110 let moves = x86_64::FRAME.moves(self.gpr).expect("a class the target says how to move");
2111 let load = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", moves.load)));
2112 self.stack.walks_frames = true;
2113
2114 // Where the walk is up to. The frame pointer to begin with, and the register the last load
2115 // wrote after that.
2116 let reg = self.new_reg(result);
2117 let mut base = mir::Reg::physical(self.conv.frame_pointer);
2118 for link in 0..depth {
2119 // The last load of a walk that is looking for a frame writes the answer itself, which
2120 // is what keeps a walk of so many links that many instructions and not one more.
2121 let ends_here = link + 1 == depth && !returning;
2122 let next = if ends_here { reg } else { self.out.new_vreg(self.gpr) };
2123 let at = mir::Mem::at(mir::Operand::read(base, self.gpr));
2124 self.out.build(block, load).at(span).def(next, self.gpr).mem(at).finish();
2125 base = next;
2126 }
2127
2128 if returning {
2129 let up = i32::try_from(self.conv.return_address).expect("a word above the frame");
2130 let at = mir::Mem::at(mir::Operand::read(base, self.gpr)).plus(up);
2131 self.out.build(block, load).at(span).def(reg, self.gpr).mem(at).finish();
2132 } else if depth == 0 {
2133 // The one case with no load in it at all: the frame this function is running in is the
2134 // register itself, and a physical register is not one the allocator hands out, so the
2135 // answer is a copy of it.
2136 let mov = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", moves.mov)));
2137 self.out
2138 .build(block, mov)
2139 .at(span)
2140 .operand(mir::Operand::write(reg, self.gpr))
2141 .operand(mir::Operand::read(base, self.gpr))
2142 .finish();
2143 }
2144 Ok(())
2145 }
2146
2147 /// `__builtin_thread_pointer`, which is the front of the block [`Self::thread_address`] adds
2148 /// an offset to.
2149 ///
2150 /// The same one instruction, on its own this time and with nothing to add to it. A program
2151 /// writes this when what it wants is a number that is different in every thread and cheap to
2152 /// come by, rather than a variable of its own in the block, so there is no relocation here and
2153 /// no name for the link to resolve.
2154 fn thread_pointer(&mut self, inst: Inst) -> Result<(), Unsupported> {
2155 let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
2156 let block = self.at.expect("a block is being filled");
2157 let span = self.source.span(inst);
2158 let reg = self.new_reg(result);
2159 let load = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{GOT_LOAD}")));
2160 let at = mir::Mem::in_segment(Segment::Fs, 0);
2161 self.out.build(block, load).at(span).def(reg, self.gpr).mem(at).finish();
2162 Ok(())
2163 }
2164
2165 /// A conversion that converts nothing: the result is the operand under another type.
2166 ///
2167 /// `ptrtoint` and `inttoptr` at one width are the whole of this. An address on this machine is
2168 /// an integer as wide as the machine addresses, so a cast between the two changes what the
2169 /// type system calls the value and changes nothing about the value, and the register holding
2170 /// it is the register that already held it. The front end never writes either of them at any
2171 /// other width, because it widens or narrows around the cast rather than through it, so the
2172 /// two widths disagreeing here means the IR came from somewhere else and is refused rather
2173 /// than guessed at.
2174 ///
2175 /// Reading the operand first is what materializes it when it is a constant, which is the case
2176 /// that matters: a null pointer is an `inttoptr` of zero, and that zero has to reach a
2177 /// register before anything can call it an address.
2178 fn rename(&mut self, inst: Inst) -> Result<(), Unsupported> {
2179 let data = &self.source[inst];
2180 let [arg] = self.source[data.args] else { return Err(self.unsupported(inst)) };
2181 let result = data.first_result.ok_or_else(|| self.unsupported(inst))?;
2182 if !self.is_address_width(self.source[arg].ty)
2183 || !self.is_address_width(self.source[result].ty)
2184 {
2185 return Err(self.unsupported(inst));
2186 }
2187 let reg = self.reg_of(arg)?;
2188 self.regs[result.index()] = Some(reg);
2189 Ok(())
2190 }
2191
2192 /// One barrier, which on this machine is one instruction at the strongest ordering and no
2193 /// instruction at all at every other one.
2194 ///
2195 /// x86-64 is total store order, so the only reordering the machine does is a store followed by
2196 /// a load of a different address, and the only ordering that forbids that is sequential
2197 /// consistency. An acquire, a release and an acquire release fence are therefore already true
2198 /// of every program running here, and what a program wanted from writing one is that the
2199 /// compiler not move memory accesses across it. The optimizer has finished by the time this
2200 /// runs and nothing below reorders one access past another, so the constraint is already
2201 /// discharged and there is nothing to write.
2202 ///
2203 /// The strongest one is `mfence`, which is what gcc 16.2.0 writes for
2204 /// `__atomic_thread_fence(__ATOMIC_SEQ_CST)` and for `__sync_synchronize`. A locked instruction
2205 /// on the stack is faster on most parts and is what some compilers write instead; it is also a
2206 /// write to memory the program did not ask for, and the plain barrier is the one that says what
2207 /// it means.
2208 ///
2209 /// Written here by name rather than by a rule, for the same reason a `lea` of a symbol is:
2210 /// there is nothing in a barrier that a proof over bitvectors could discharge. It computes
2211 /// nothing, so there is no equality to state, and what makes it the right answer is the memory
2212 /// model, which the rule language cannot talk about.
2213 fn barrier(&mut self, inst: Inst) -> Result<(), Unsupported> {
2214 let Extra::Order(order) = self.source[inst].extra else {
2215 return Err(self.unsupported(inst));
2216 };
2217 if order != MemOrder::SeqCst {
2218 return Ok(());
2219 }
2220 let block = self.at.expect("a block is being filled");
2221 let span = self.source.span(inst);
2222 let fence = mir::Opcode::new(self.names.intern("x64.mfence"));
2223 self.out.build(block, fence).at(span).finish();
2224 Ok(())
2225 }
2226
2227 /// The instruction a program stops on, which is one byte pair and no operands.
2228 ///
2229 /// `ud2` is an opcode the manual promises will never be given a meaning, so a processor that
2230 /// reaches it raises the fault for an instruction it does not know, and on Linux that arrives
2231 /// at the program as `SIGILL`. That is what `__builtin_trap` is for: a stop that cannot be
2232 /// caught by anything the program installed for an ordinary error, cannot be returned from,
2233 /// and leaves the address of the fault in the core file.
2234 ///
2235 /// Why not a call to `abort`. It is two bytes against a call and a relocation, it needs no
2236 /// library, and it works in the places this one is written most, which are a kernel and a
2237 /// freestanding program that has no `abort` to call. gcc 16.2.0 writes `ud2` here too.
2238 fn trap(&mut self, inst: Inst) {
2239 let block = self.at.expect("a block is being filled");
2240 let span = self.source.span(inst);
2241 let stop = mir::Opcode::new(self.names.intern("x64.ud2"));
2242 self.out.build(block, stop).at(span).finish();
2243 }
2244
2245 /// One hint that an address is about to be used, which is one instruction and no promise.
2246 ///
2247 /// Four instructions on this machine and the locality picks between them, which is what the
2248 /// number means: how much of the data will still be wanted after the access. None of it wanted
2249 /// is `prefetchnta`, which brings the line in without keeping it, and all of it wanted is
2250 /// `prefetcht0`, which brings it as close as the machine can. The two in between are the levels
2251 /// between those. Measured against gcc 16.2.0 on x86-64 rather than read off the manual: zero
2252 /// gives `prefetchnta`, one `prefetcht2`, two `prefetcht1` and three `prefetcht0`.
2253 ///
2254 /// Whether the access will write is not read here, and that is this machine rather than an
2255 /// omission. The write hint is `prefetchw`, which is not in the base instruction set, and gcc
2256 /// writes it only when the command line said the part has it. So a prefetch for a write is the
2257 /// same instruction as a prefetch for a read, which is what gcc 16.2.0 writes without
2258 /// `-mprfchw`, and the difference is carried in the IR for a target that can use it.
2259 ///
2260 /// The address goes in the addressing mode rather than in an operand, the way a store's does.
2261 /// It is built here as the plainest one there is, a register and nothing else, because what
2262 /// arrives is a value and folding an addition into the mode is a rule's job and no rule reaches
2263 /// this instruction. An address the program computed is therefore one `lea` or one add in front
2264 /// of this, which is what it would have been for the load the hint is about anyway.
2265 fn hint(&mut self, inst: Inst) -> Result<(), Unsupported> {
2266 let Extra::Prefetch(hint) = self.source[inst].extra else {
2267 return Err(self.unsupported(inst));
2268 };
2269 let args: Vec<Value> = self.source[self.source[inst].args].to_vec();
2270 let [address] = args[..] else { return Err(self.unsupported(inst)) };
2271 let name = match hint.locality {
2272 0 => "prefetch_nta",
2273 1 => "prefetch_t2",
2274 2 => "prefetch_t1",
2275 PrefetchHint::MOST => "prefetch_t0",
2276 // Nothing else exists. The checker reads a locality outside the range as zero and the
2277 // verifier refuses one that got here another way, so this is a hint that was built
2278 // rather than checked, and the safe answer for a hint is to write no instruction.
2279 _ => return Err(self.unsupported(inst)),
2280 };
2281 let base = self.reg_of(address)?;
2282 let block = self.at.expect("a block is being filled");
2283 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
2284 self.out
2285 .build(block, opcode)
2286 .at(self.source.span(inst))
2287 .mem(mir::Mem::at(mir::Operand::read(base, self.gpr)))
2288 .finish();
2289 Ok(())
2290 }
2291
2292 /// One compare and exchange, which is the instruction every other atomic on this machine is
2293 /// built out of.
2294 ///
2295 /// What the IR asks for is: read what is at an address, compare it against a value the program
2296 /// expected, put a second value there if the two were equal, and say both what was read and
2297 /// whether the exchange happened. The machine has exactly that instruction, and the `lock` in
2298 /// front of it is what makes the whole of it one step as far as every other processor is
2299 /// concerned.
2300 ///
2301 /// The ordering is not read here, and that is the memory model rather than an omission. A
2302 /// locked instruction on x86-64 is a full barrier whatever the program asked for, so a relaxed
2303 /// compare and exchange and a sequentially consistent one are the same instruction, and there
2304 /// is nothing weaker to emit for the weaker orderings. The failure ordering is not read for the
2305 /// same reason.
2306 ///
2307 /// The two values it produces are why this is written by name. The one the program compares
2308 /// against and the one it gets back are both `rax`, which the instruction reads and writes
2309 /// without being told, and the table says so with a fixed constraint at each end rather than
2310 /// leaving the allocator to find out. The second value is the byte behind it, which is the zero
2311 /// flag read out by a `sete`, and it is a definition of the same instruction so that the
2312 /// allocator knows the two are live together and never gives the byte the register the answer
2313 /// is in.
2314 fn exchange(&mut self, inst: Inst) -> Result<(), Unsupported> {
2315 let args: Vec<Value> = self.source[self.source[inst].args].to_vec();
2316 let results: Vec<Value> = self.source[inst].results().collect();
2317 let [addr, expected, desired] = args[..] else { return Err(self.unsupported(inst)) };
2318 let [old, exchanged] = results[..] else { return Err(self.unsupported(inst)) };
2319
2320 // A value the machine can compare in one instruction, which is an integer or an address at
2321 // one of the four widths it has a compare and exchange for. Anything else is a type this
2322 // has no instruction for rather than a program that is wrong, and the front end refuses it
2323 // before ever getting here.
2324 let ty = self.source[old].ty;
2325 let bits = if ty.is_ptr() { ADDRESS_BITS } else { ty.bits() };
2326 if (!ty.is_int() && !ty.is_ptr()) || !matches!(bits, 8 | 16 | 32 | 64) {
2327 return Err(self.unsupported(inst));
2328 }
2329
2330 let base = self.reg_of(addr)?;
2331 let want = self.reg_of(expected)?;
2332 let put = self.reg_of(desired)?;
2333 let got = self.new_reg(old);
2334 let flag = self.new_reg(exchanged);
2335
2336 let name = format!("cmpxchg_{bits}");
2337 let form = x86_64::form(&name).ok_or_else(|| self.unsupported(inst))?;
2338 let block = self.at.expect("a block is being filled");
2339 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
2340 let mut build = self.out.build(block, opcode).at(self.source.span(inst));
2341 for (desc, reg) in form.operands().iter().zip([got, flag, want, put]) {
2342 let operand = mir::Operand {
2343 reg,
2344 class: desc.class,
2345 role: desc.role,
2346 constraint: desc.constraint,
2347 };
2348 build = build.operand(operand);
2349 }
2350 build.mem(mir::Mem::at(mir::Operand::read(base, self.gpr))).finish();
2351 Ok(())
2352 }
2353
2354 /// One read modify write, for the three operations this machine does in a single instruction.
2355 ///
2356 /// What the IR asks for is: read what is at an address, do something to it, put the answer back,
2357 /// say what was there before, and let nothing get between the three steps. The machine has
2358 /// `xchg` for putting a value there and `lock xadd` for adding one, and both leave what they
2359 /// found in the register the operand arrived in, which is why the value that comes back and the
2360 /// value that went in are one register here.
2361 ///
2362 /// A subtraction is the add over the negated operand, which is right at every width because the
2363 /// machine's arithmetic wraps and negating then adding is subtracting in two's complement
2364 /// whatever the operands were. The negate is a separate instruction in front, over a register of
2365 /// its own, so that the value the program handed over is not the one written on: an operand may
2366 /// be live after this and a program that read it again would read the negation.
2367 ///
2368 /// The ordering is not read, for the reason the compare and exchange beside this does not read
2369 /// it. `xchg` with memory locks the bus whether it is asked to or not and `lock xadd` is asked
2370 /// to, so both are full barriers on this machine and there is nothing weaker to fall to.
2371 ///
2372 /// Eight of the other ten never arrive, because `crate::retry` turned each of them into a loop
2373 /// around a compare and exchange before anything here saw it. The two that do arrive are the
2374 /// ones on floating values, and they are refused: a compare and exchange of a float wants the
2375 /// value carried through an integer of the same width, and an eighty bit float has no such
2376 /// width. Neither family of builtins can write one yet either, so a program that reaches this
2377 /// refusal is a program that reached an unimplemented builtin first.
2378 fn modify(&mut self, inst: Inst) -> Result<(), Unsupported> {
2379 let Extra::Rmw(op, _) = self.source[inst].extra else {
2380 return Err(self.unsupported(inst));
2381 };
2382 let args: Vec<Value> = self.source[self.source[inst].args].to_vec();
2383 let [addr, operand] = args[..] else { return Err(self.unsupported(inst)) };
2384 let old = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
2385
2386 // A value the machine can exchange in one instruction, which is an integer at one of the
2387 // four widths it has these for. A pointer arrives as an address, so it is an integer by the
2388 // time it is here, and anything else is a type this has no instruction for.
2389 let ty = self.source[old].ty;
2390 if !ty.is_int() || !matches!(ty.bits(), 8 | 16 | 32 | 64) {
2391 return Err(self.unsupported(inst));
2392 }
2393 let name = match op {
2394 RmwOp::Xchg => format!("xchg_{}", ty.bits()),
2395 RmwOp::Add | RmwOp::Sub => format!("xadd_{}", ty.bits()),
2396 _ => return Err(self.unsupported(inst)),
2397 };
2398
2399 let base = self.reg_of(addr)?;
2400 let mut put = self.reg_of(operand)?;
2401 let block = self.at.expect("a block is being filled");
2402 let span = self.source.span(inst);
2403 if op == RmwOp::Sub {
2404 let negated = self.out.new_vreg(self.gpr);
2405 let negate =
2406 mir::Opcode::new(self.names.intern(&format!("{PREFIX}neg_r_{}", ty.bits())));
2407 let form = x86_64::form(&format!("neg_r_{}", ty.bits()))
2408 .ok_or_else(|| self.unsupported(inst))?;
2409 let mut build = self.out.build(block, negate).at(span);
2410 for (desc, reg) in form.operands().iter().zip([negated, put]) {
2411 build = build.operand(mir::Operand {
2412 reg,
2413 class: desc.class,
2414 role: desc.role,
2415 constraint: desc.constraint,
2416 });
2417 }
2418 build.finish();
2419 put = negated;
2420 }
2421
2422 let got = self.new_reg(old);
2423 let form = x86_64::form(&name).ok_or_else(|| self.unsupported(inst))?;
2424 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{name}")));
2425 let mut build = self.out.build(block, opcode).at(span);
2426 for (desc, reg) in form.operands().iter().zip([got, put]) {
2427 build = build.operand(mir::Operand {
2428 reg,
2429 class: desc.class,
2430 role: desc.role,
2431 constraint: desc.constraint,
2432 });
2433 }
2434 build.mem(mir::Mem::at(mir::Operand::read(base, self.gpr))).finish();
2435 Ok(())
2436 }
2437
2438 /// One `asm` statement.
2439 ///
2440 /// An empty template is most of the inline assembly in a test suite, and it is not a corner
2441 /// case somebody wrote by accident. A program that wants a value computed where it stands, or a
2442 /// loop the optimizer must not touch, writes `asm volatile ("" : : : "memory")`, and forty
2443 /// years of bug reports about optimizers are full of them. What such a statement asks for is
2444 /// the barrier and the operand places, and no instructions at all.
2445 ///
2446 /// So the operands are the half that is always real: a constraint says where a value has to be,
2447 /// and where it has to be is still true when the template between them is empty.
2448 ///
2449 /// What the constraints ask for, on an empty template, is only ever that two operands share a
2450 /// place. Nothing reads a register no text names, so `"r"` on its own asks for a register and
2451 /// no particular one, and any register at all answers it. A matching constraint is different,
2452 /// because it says the output the assembly leaves is the place the input arrived in, and with
2453 /// no instructions between them that is the input unchanged. So it is a rename and not a move:
2454 /// the value is already in a register and the result is that register.
2455 ///
2456 /// An output nothing is tied to and no instruction writes is whatever the assembly left there,
2457 /// which for a template that writes nothing is whatever was in the register. That is a value
2458 /// the program is not entitled to, and this writes a zero rather than reading one, because the
2459 /// allocator has to be given a definition before a use whatever the program is entitled to.
2460 ///
2461 /// # A template with instructions in it
2462 ///
2463 /// [`x86_64::read`] turns the text into the opcodes this backend already has, which is what
2464 /// `spec/11-asm-objects-debug.md` section 11.1 asks for: the machine is described once, and an
2465 /// instruction a program wrote is looked up in that description rather than copied through to
2466 /// an assembler that has one of its own. So nothing here assembles anything. What it does is
2467 /// put the statement's operands where the opcode holds them, and from there an `asm` statement
2468 /// is ordinary machine code: the allocator picks the registers, the listing and the object file
2469 /// are written from the same table as every other instruction, and a spill around one works
2470 /// because there is nothing left about it for a spill to get wrong.
2471 ///
2472 /// Three things are refused, all for one reason, which is that placing them by a guess gives a
2473 /// program that assembles into something other than what it says.
2474 ///
2475 /// A register the template named itself. The registers an instruction here names are the ones
2476 /// the allocator handed out, and a name in the text is a claim on a register nobody told the
2477 /// allocator about. A register a constraint letter names is a different thing and is placed,
2478 /// which the paragraph below is about: there the statement said which of its own operands is
2479 /// in the register, and a name in the middle of a template says no such thing.
2480 ///
2481 /// An output the template writes more than once, which is one place with two definitions in it,
2482 /// and the machine IR between here and the allocator has one definition per register by
2483 /// construction. An output tied to an input and written once is not that: it is two registers
2484 /// the description ties together, which is what [`Place`] is about.
2485 ///
2486 /// An operand read where the opcode writes, or written where it reads. An output that has not
2487 /// been written yet is not a value, and an input the assembly writes over is a value something
2488 /// else may still be using.
2489 ///
2490 /// # A register the instruction uses without being told
2491 ///
2492 /// An instruction may reach a register its text does not name, and `cpuid` is all of them at
2493 /// once: the leaf goes in `eax`, the subleaf in `ecx`, and the answer comes back in all four
2494 /// registers. The description holds every bit of that already, so what is left is to say which
2495 /// of the statement's operands is in each of those registers, and the constraint letter is the
2496 /// one thing in an assembly statement that says it. `"=a"` is an output in `rax` and `"c"` is
2497 /// an input in `rcx`, which is why a program writing `cpuid` writes its constraints that way
2498 /// and has no choice about it.
2499 ///
2500 /// A register no letter named is one the statement put nothing in, and that is the usual case
2501 /// rather than an unusual one, since an instruction that answers four questions is written by
2502 /// programs that asked one. A write of one is the register being destroyed and gets a register
2503 /// of its own, which is what tells the allocator to keep everything else out of it. A read of
2504 /// one is a register the instruction looks at and the program never filled, which gets a zero
2505 /// for the reason [`Self::undefined`] gives.
2506 ///
2507 /// # The clobber list
2508 ///
2509 /// Read now, as the registers it names being written by every instruction of the template. By
2510 /// every one rather than by one of them, because the list says the assembly as a whole leaves
2511 /// them ruined and nothing here knows which line did it. Every entry has to be a register this
2512 /// machine has a name for or the statement is refused, since a name nobody read is a register
2513 /// nobody is keeping out of.
2514 ///
2515 /// `memory` and `cc` are the two entries that are not registers and both are skipped. `memory`
2516 /// says the assembly touches storage, which is already true of every `asm` this writes and is
2517 /// nothing a register list could hold. `cc` says it ruins the condition flags, and the flag
2518 /// tracking already has that from the instructions the template was read into, since it takes
2519 /// every instruction it does not recognize as writing them and every instruction here is one
2520 /// this machine describes.
2521 ///
2522 /// A clobber the instruction already writes is left off it. `cpuid` writes all four registers
2523 /// by description, and a statement listing three of them as clobbers as well is saying the
2524 /// same thing twice, which the allocator would read as one register with two definitions.
2525 ///
2526 /// On a template with nothing in it the list is ignored, as it was before, since a template
2527 /// with no instructions ruins nothing whatever it said about what it ruins.
2528 fn assembly(&mut self, inst: Inst) -> Result<(), Unsupported> {
2529 let data = &self.source[inst];
2530 let Extra::Asm(asm) = data.extra else { return Err(self.unsupported(inst)) };
2531 let info = self.source[asm];
2532 if !self.source[info.targets].is_empty() {
2533 return Err(Unsupported::Assembly { inst, refused: Written::Goto });
2534 }
2535 let refused = || Unsupported::Assembly { inst, refused: Written::Operand };
2536
2537 let constraints = self.names.resolve(info.constraints).to_string();
2538 let results: Vec<Value> = data.results().collect();
2539 let operands = AsmOperands::read(&constraints, &results, &self.source[data.args])
2540 .ok_or_else(refused)?;
2541 let list: Vec<AsmOperand> = operands.iter().copied().collect();
2542
2543 // Read after the constraints and not before them, because a mnemonic whose suffix the
2544 // program left off is read at the width of the operands it names, and the operands are
2545 // what the constraints are a list of.
2546 let widths: Vec<Option<x86_64::Width>> = list
2547 .iter()
2548 .map(|operand| {
2549 let ty = self.source[operand.result.or(operand.value)?].ty;
2550 if !ty.is_scalar() {
2551 return None;
2552 }
2553 x86_64::Width::of_bits(if ty.is_ptr() { ADDRESS_BITS } else { ty.bits() })
2554 })
2555 .collect();
2556 let template = self.names.resolve(info.template).to_string();
2557 let lines = if template.trim().is_empty() {
2558 Vec::new()
2559 } else {
2560 x86_64::read(&template, &widths)
2561 .ok_or(Unsupported::Assembly { inst, refused: Written::Template })?
2562 };
2563
2564 // Which operands the template writes, counted before anything is placed, because the answer
2565 // decides where each of the three below comes from and one instruction may name an operand
2566 // that a later one writes.
2567 let mut writes = vec![0usize; list.len()];
2568 for line in &lines {
2569 let form = x86_64::form(line.opcode).ok_or_else(refused)?;
2570 for (desc, piece) in form.operands().iter().zip(&line.operands) {
2571 // An operand the instruction reaches without its text saying so is the statement's
2572 // only when a constraint letter put something there. One that is nobody's writes
2573 // nothing of the program's, so it is counted nowhere and is dealt with where it is
2574 // placed.
2575 let index = match *piece {
2576 x86_64::Piece::Operand { index, .. } => index,
2577 x86_64::Piece::Implicit { reg } => match bound(&list, reg, desc.role) {
2578 Some(index) => index,
2579 None => continue,
2580 },
2581 x86_64::Piece::Reg { .. } => continue,
2582 };
2583 if matches!(desc.role, Role::Def | Role::EarlyDef) {
2584 *writes.get_mut(index).ok_or_else(refused)? += 1;
2585 }
2586 }
2587 }
2588
2589 // Where every operand is. Worked out in full before the first instruction is written, since
2590 // reading a value may be what puts it in a register in the first place, and that has to
2591 // happen in front of the assembly rather than in the middle of it.
2592 let mut places: Vec<Place> = vec![Place::default(); list.len()];
2593 for (index, operand) in list.iter().copied().enumerate() {
2594 let Some(result) = operand.result else {
2595 // An input, or an output the assembly was handed the address of, and both are a
2596 // value that arrives in a register and is read out of it.
2597 places[index].read = Some(self.reg_of(operand.value.ok_or_else(refused)?)?);
2598 continue;
2599 };
2600 let ty = self.source[result].ty;
2601 if on_x87(ty) || writes[index] > 1 {
2602 return Err(refused());
2603 }
2604 let tied = operands.tied_to(index);
2605 if let Some(from) = tied {
2606 if self.class_of(self.source[from].ty) != self.class_of(ty) {
2607 return Err(refused());
2608 }
2609 places[index].read = Some(self.reg_of(from)?);
2610 }
2611 if writes[index] == 1 {
2612 places[index].write = Some(self.new_reg(result));
2613 continue;
2614 }
2615 match tied {
2616 // The place the input arrived in, which the assembly wrote nothing over. One
2617 // register, so this is a rename rather than a move.
2618 Some(_) => {
2619 let reg = places[index].read.ok_or_else(refused)?;
2620 self.regs[result.index()] = Some(reg);
2621 places[index].write = Some(reg);
2622 }
2623 None => {
2624 self.undefined(inst, result)?;
2625 places[index].write = self.regs[result.index()];
2626 }
2627 }
2628 }
2629
2630 // Worked out once for the whole template, since the list is one list and every instruction
2631 // of the template gets it. Not worked out at all for a template with no instructions, which
2632 // is where there is nothing for it to go on.
2633 let clobbers = self.names.resolve(info.clobbers).to_string();
2634 let clobbered =
2635 if lines.is_empty() { Vec::new() } else { Self::clobbered(inst, &clobbers)? };
2636
2637 for line in &lines {
2638 self.instruction(inst, line, &places, &list, &clobbered)?;
2639 }
2640 Ok(())
2641 }
2642
2643 /// The registers a clobber list names, in the order it named them.
2644 ///
2645 /// Nothing is dropped. A name this has no register for is refused, because the list is the
2646 /// program telling the compiler which registers it may not leave anything in, and an entry
2647 /// nobody read is a register something may still be left in. See [`Self::assembly`] for the
2648 /// two entries that are not registers and for why they are skipped rather than refused.
2649 fn clobbered(inst: Inst, clobbers: &str) -> Result<Vec<PhysReg>, Unsupported> {
2650 let refused = || Unsupported::Assembly { inst, refused: Written::Clobber };
2651 let mut named = Vec::new();
2652 for entry in clobbers.split(',') {
2653 let entry = entry.trim().trim_matches('"');
2654 // The sigil is optional in a clobber list and means nothing when it is there, unlike
2655 // in a template, where it is what tells a register from an operand.
2656 let entry = entry.strip_prefix('%').unwrap_or(entry);
2657 if entry.is_empty() || entry == "memory" || entry == "cc" {
2658 continue;
2659 }
2660 let (reg, _) = x86_64::gpr_named(entry).ok_or_else(refused)?;
2661 if !named.contains(®) {
2662 named.push(reg);
2663 }
2664 }
2665 Ok(named)
2666 }
2667
2668 /// One instruction of a template, as the machine instruction it was read back into.
2669 fn instruction(
2670 &mut self,
2671 inst: Inst,
2672 line: &x86_64::Line,
2673 places: &[Place],
2674 list: &[AsmOperand],
2675 clobbered: &[PhysReg],
2676 ) -> Result<(), Unsupported> {
2677 let refused = || Unsupported::Assembly { inst, refused: Written::Operand };
2678 let form = x86_64::form(line.opcode).ok_or_else(refused)?;
2679 let mut built = Vec::with_capacity(line.operands.len() + clobbered.len());
2680 for (desc, piece) in form.operands().iter().zip(&line.operands) {
2681 built.push(self.placed(inst, *desc, *piece, places, list)?);
2682 }
2683 // The clobbers go in among the definitions rather than behind the reads, because an operand
2684 // vector in the machine IR is every definition and then every use and what counts them
2685 // reads that order rather than each operand's role.
2686 let defs = built.iter().take_while(|operand| operand.role.is_def()).count();
2687 let mut added = 0usize;
2688 for ® in clobbered {
2689 if form.operands().iter().any(|desc| desc.constraint == Constraint::Fixed(reg)) {
2690 continue;
2691 }
2692 built.insert(defs, mir::Operand::write(mir::Reg::physical(reg), self.gpr));
2693 added += 1;
2694 }
2695 // A constraint tying one operand to another names it by its place in this vector, and the
2696 // clobbers were put in the middle of the vector, so everything behind them moved. The
2697 // description is written against an instruction with no clobbers in it and cannot know
2698 // that, which makes this the one place the two numberings have to be reconciled.
2699 for operand in &mut built {
2700 if let Constraint::Reuse(at) = operand.constraint {
2701 if usize::from(at) >= defs {
2702 let moved = usize::from(at) + added;
2703 operand.constraint =
2704 Constraint::Reuse(u8::try_from(moved).map_err(|_| refused())?);
2705 }
2706 }
2707 }
2708 let at = match line.at {
2709 Some(at) => Some(self.addressed(inst, at, places)?),
2710 None => None,
2711 };
2712
2713 let block = self.at.expect("a block is being filled");
2714 let span = self.source.span(inst);
2715 let opcode = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", line.opcode)));
2716 let mut build = self.out.build(block, opcode).at(span);
2717 for operand in built {
2718 build = build.operand(operand);
2719 }
2720 if let Some(value) = line.imm {
2721 build = build.imm(value);
2722 }
2723 if let Some(mem) = at {
2724 build = build.mem(mem);
2725 }
2726 build.finish();
2727 Ok(())
2728 }
2729
2730 /// One operand of one instruction of a template, in the register the statement put it in.
2731 fn placed(
2732 &mut self,
2733 inst: Inst,
2734 desc: OperandDesc,
2735 piece: x86_64::Piece,
2736 places: &[Place],
2737 list: &[AsmOperand],
2738 ) -> Result<mir::Operand, Unsupported> {
2739 let refused = || Unsupported::Assembly { inst, refused: Written::Operand };
2740 // A register the instruction reaches without its text naming it belongs to whichever of the
2741 // statement's operands a constraint letter put there, and to nobody when no letter did.
2742 // There is no width to check in that case: the operand is the register the letter named and
2743 // the instruction does what it does to it, which is what a program writing `"=a"` asked for.
2744 let (index, width) = match piece {
2745 x86_64::Piece::Operand { index, width } => (index, Some(width)),
2746 x86_64::Piece::Implicit { reg } => match bound(list, reg, desc.role) {
2747 Some(index) => (index, None),
2748 None => return self.spare(inst, desc),
2749 },
2750 x86_64::Piece::Reg { .. } => return Err(refused()),
2751 };
2752 let operand = list.get(index).copied().ok_or_else(refused)?;
2753 // The two halves of an operand written `+`, which arrives in one register and leaves in
2754 // another with the allocator told to make them the same one. Everything else has one of
2755 // the two and asking for the other is the refusal below.
2756 let place = places.get(index).copied().ok_or_else(refused)?;
2757 let reg = match desc.role {
2758 Role::Use => place.read,
2759 Role::Def | Role::EarlyDef => place.write,
2760 }
2761 .ok_or_else(refused)?;
2762
2763 // Read where the opcode reads and written where it writes, which is what the first half of
2764 // this asks. An output has a result and an input has a value, and an output written `+` has
2765 // both, because it is read before it is written.
2766 let placeable = match desc.role {
2767 Role::Use => operand.value.is_some(),
2768 Role::Def | Role::EarlyDef => operand.result.is_some(),
2769 };
2770 let ty = match (operand.result, operand.value) {
2771 (Some(result), _) => self.source[result].ty,
2772 (None, Some(value)) => self.source[value].ty,
2773 (None, None) => return Err(refused()),
2774 };
2775 let bits = if ty.is_ptr() { ADDRESS_BITS } else { ty.bits() };
2776 if !placeable || self.class_of(ty) != desc.class {
2777 return Err(refused());
2778 }
2779 if width.is_some_and(|width| bits != width.bits()) {
2780 return Err(refused());
2781 }
2782 Ok(mir::Operand { reg, class: desc.class, role: desc.role, constraint: desc.constraint })
2783 }
2784
2785 /// A register an instruction of a template uses and the statement put nothing in.
2786 ///
2787 /// A write of one is the register being destroyed, which is what a clobber list is usually
2788 /// written to say and what an instruction with more answers than the program asked for does
2789 /// anyway: `cpuid` writes all four registers whether or not the statement wanted all four. A
2790 /// register of its own is the whole of what that needs, since a value nothing reads is one the
2791 /// allocator may put anywhere and is told about so that nothing else is put there.
2792 ///
2793 /// A read of one is a register the instruction looks at and the program never filled, which
2794 /// gcc leaves as whatever happened to be there. A zero is written instead, for the reason
2795 /// [`Self::undefined`] gives: the allocator has to be given a definition before a use, and a
2796 /// zero is the one answer that reads the same on every run.
2797 fn spare(&mut self, inst: Inst, desc: OperandDesc) -> Result<mir::Operand, Unsupported> {
2798 let refused = Unsupported::Assembly { inst, refused: Written::Operand };
2799 if desc.class != self.gpr {
2800 return Err(refused);
2801 }
2802 let reg = self.out.new_vreg(desc.class);
2803 if !desc.role.is_def() {
2804 let block = self.at.expect("a block is being filled");
2805 let span = self.source.span(inst);
2806 let put = mir::Opcode::new(self.names.intern(&format!("{PREFIX}mov_ri_64")));
2807 self.out.build(block, put).at(span).def(reg, desc.class).imm(0).finish();
2808 }
2809 Ok(mir::Operand { reg, class: desc.class, role: desc.role, constraint: desc.constraint })
2810 }
2811
2812 /// The address one instruction of a template reads or writes.
2813 fn addressed(
2814 &mut self,
2815 inst: Inst,
2816 at: x86_64::At,
2817 places: &[Place],
2818 ) -> Result<mir::Mem, Unsupported> {
2819 let refused = || Unsupported::Assembly { inst, refused: Written::Operand };
2820 let base = match at.base {
2821 None => None,
2822 Some(x86_64::Piece::Operand { index, .. }) => {
2823 // The register an address is counted from is read and never written, whatever the
2824 // instruction does to what it finds there.
2825 let reg = places.get(index).and_then(|place| place.read).ok_or_else(refused)?;
2826 Some(mir::Operand::read(reg, self.gpr))
2827 }
2828 // An address counted from a register the instruction reaches without being told is
2829 // not something this machine has: every addressing mode is written out in the text it
2830 // is part of, so a base that got here another way is a base nothing wrote down.
2831 Some(x86_64::Piece::Reg { .. } | x86_64::Piece::Implicit { .. }) => {
2832 return Err(refused());
2833 }
2834 };
2835 Ok(mir::Mem { base, scale: 1, disp: at.disp, segment: at.segment, ..mir::Mem::default() })
2836 }
2837
2838 /// A register holding a value the program has no claim on, written as a zero.
2839 ///
2840 /// Every other way of saying it costs the same instruction or needs a word the machine IR does
2841 /// not have, and a zero is the one that reads the same on every run.
2842 fn undefined(&mut self, inst: Inst, result: Value) -> Result<(), Unsupported> {
2843 let ty = self.source[result].ty;
2844 let refused = Unsupported::Assembly { inst, refused: Written::Operand };
2845 if self.class_of(ty) != self.gpr || !matches!(ty.bits(), 8 | 16 | 32 | 64) {
2846 return Err(refused);
2847 }
2848 let block = self.at.expect("a block is being filled");
2849 let span = self.source.span(inst);
2850 let reg = self.new_reg(result);
2851 let put = mir::Opcode::new(self.names.intern(&format!("{PREFIX}mov_ri_{}", ty.bits())));
2852 self.out.build(block, put).at(span).def(reg, self.gpr).imm(0).finish();
2853 Ok(())
2854 }
2855
2856 /// Whether a type is the width an address is, which is what makes a cast to or from one free.
2857 fn is_address_width(&self, ty: Type) -> bool {
2858 ty.is_ptr() || (ty.is_int() && ty.bits() == ADDRESS_BITS)
2859 }
2860
2861 /// Where a block goes, which in machine IR is on the block rather than on its terminator.
2862 ///
2863 /// That is why no rule ever names a block: a branch is selected for what it reads and the
2864 /// edges are copied across here, arguments and all. The arguments are read last, after every
2865 /// instruction of the block is written, because an argument that is a constant is
2866 /// materialized where it is first wanted and the end of the block is where an edge wants it.
2867 ///
2868 /// Which is not quite the end. A block that leaves two ways has the branch as its last
2869 /// instruction, and a block that leaves through a register has the indirect jump as its last,
2870 /// and anything appended after either is something it has already jumped past, so a constant
2871 /// materialized here would be a register the block below reads and nothing ever writes. The
2872 /// one that was there is put back on the end when that happened, which is the only reordering
2873 /// anything in this crate does and is why it is remembered before a single argument is read.
2874 fn edges(&mut self, block: Block, out: mir::Block) -> Result<(), Unsupported> {
2875 let Some(term) = self.source.terminator(block) else { return Ok(()) };
2876 let leaves = matches!(self.source[term].opcode, Opcode::BrIf | Opcode::IndirectBr);
2877 let branch = if leaves { self.out.terminator(out) } else { None };
2878
2879 let calls: Vec<rucc_ir::BlockCall> = self.source.successors(term).collect();
2880 let mut succs = Vec::with_capacity(calls.len());
2881 for call in calls {
2882 let args: Vec<Value> = self.source[call.args].to_vec();
2883 let mut regs = Vec::with_capacity(args.len());
2884 for value in args {
2885 // The address of where the value is rather than the value, for the one type a
2886 // register holds none of. The block on the other side copies the bytes out of it
2887 // into a slot of its own, which is what makes a second edge into the same block
2888 // safe.
2889 let reg = if on_x87(self.source[value].ty) {
2890 self.x87_slot(value)
2891 } else {
2892 self.reg_of(value)?
2893 };
2894 regs.push(reg);
2895 }
2896 succs.push(mir::BlockCall::with(self.out_block(call.block), regs));
2897 }
2898 if let Some(branch) = branch {
2899 if self.out.terminator(out) != Some(branch) {
2900 self.out.remove_inst(branch);
2901 self.out.append_inst(out, branch);
2902 }
2903 }
2904 *self.out.succs_mut(out) = succs;
2905 Ok(())
2906 }
2907
2908 /// The machine IR block an IR block became.
2909 fn out_block(&self, block: Block) -> mir::Block {
2910 self.blocks[block.index()].expect("every block was created before any was filled")
2911 }
2912
2913 /// The parameters of the entry block, which are the function's arguments.
2914 ///
2915 /// They are not block parameters in the machine IR and they cannot be. A block parameter is
2916 /// given its value by a move on the edge into the block, and there is no edge into an entry
2917 /// block, so what arrives in a function is the convention's to say. [`crate::abi`] is what
2918 /// says it.
2919 ///
2920 /// The ones past the last register arrived in the caller's memory and are read out of it, and
2921 /// the loads that read them come back here so that the frame can finish them the way it
2922 /// finishes an `alloca`.
2923 fn arrive(&mut self, block: Block, out: mir::Block) -> Result<(), Unsupported> {
2924 let params = self.source[block].params.clone();
2925 // The type of each is the block's answer and what the ABI asks of it is the signature's,
2926 // and the two lists are the same list: a parameter the classification turned into a
2927 // pointer is a pointer in the block too. A block with more parameters than the signature
2928 // names is not one the front end writes, and each of those is taken as a plain value.
2929 let asked: Vec<Abi> = self.source.signature().params.iter().map(|it| it.abi).collect();
2930 let types: Vec<Param> = params
2931 .iter()
2932 .enumerate()
2933 .map(|(index, &value)| {
2934 let abi = asked.get(index).copied().unwrap_or_default();
2935 Param { ty: self.source[value].ty, abi }
2936 })
2937 .collect();
2938 // A save area for a function that takes arguments its signature does not name, on a
2939 // convention whose list is the four field one. Windows is the other kind and has no area at
2940 // all, so a `va_start` in one is refused rather than built wrong.
2941 let variadic = self.source.signature().variadic && !self.conv.shared_positions;
2942 let area = variadic.then(|| varargs::Area::of(self.conv));
2943 let arrived = abi::entry(&mut self.out, out, &types, self.conv, self.names, area)
2944 .map_err(|(index, missing)| Unsupported::Argument { index, missing })?;
2945 for (¶m, reg) in params.iter().zip(&arrived.regs) {
2946 self.regs[param.index()] = Some(*reg);
2947 }
2948 if let Some(area) = area {
2949 self.save_area(out, &arrived, area);
2950 }
2951 self.stack.arguments.extend(arrived.stack);
2952 Ok(())
2953 }
2954
2955 /// The prologue of a variadic function, which is every argument register it was handed written
2956 /// into the frame.
2957 ///
2958 /// Every one the signature did not name, that is. Which of those hold anything is a thing only
2959 /// the caller knew and there is nothing here to ask, so all of them are written, and the ones a
2960 /// named parameter took are not, because `va_start` sets the two offsets past them and nothing
2961 /// ever reads their slots.
2962 ///
2963 /// What that costs is up to fourteen stores in the prologue of a function that may read none of
2964 /// them, and the convention's answer to that is the count of vector registers in `%al`, which
2965 /// lets a callee skip the eight vector stores when the call passed no floats. Skipping them is a
2966 /// branch in a prologue, and a prologue is written long after this by [`crate::finish`], which
2967 /// has no blocks to branch between. So they are all written every time, which is correct and is
2968 /// what `-O0` costs. Issue #323 is the branch.
2969 ///
2970 /// A vector register is written all sixteen bytes at a time, because a `_Float128` fills one and
2971 /// a `va_arg` of a quad reads the slot back whole. gcc writes the same sixteen with the same
2972 /// instruction, which is what [`crate::varargs`] says a list has to be built out of.
2973 ///
2974 /// The address is computed once into a register rather than written as a displacement off the
2975 /// stack pointer, because a displacement into a frame is not known until after allocation and
2976 /// one `lea` costs less than a fixup list for a dozen stores. It is the same `lea` an `alloca`
2977 /// gets and [`crate::finish`] fills it in the same way.
2978 fn save_area(&mut self, out: mir::Block, arrived: &abi::Arrived, area: varargs::Area) {
2979 let save = self.stack.locals.len();
2980 self.stack.locals.push(Local { size: area.size, align: varargs::VECTOR_SLOT });
2981 self.varargs = Some(Varargs {
2982 save,
2983 incoming: arrived.used,
2984 integers: u32::try_from(arrived.took.0).unwrap_or(0) * area.stride(false),
2985 floats: area.starts_at(true)
2986 + u32::try_from(arrived.took.1).unwrap_or(0) * area.stride(true),
2987 });
2988
2989 let base = self.frame_address(out, save);
2990 for &(reg, class, at) in &arrived.spare {
2991 let name = if class == self.gpr { "x64.mov_mr_64" } else { "x64.movaps_mr" };
2992 let store = mir::Opcode::new(self.names.intern(name));
2993 let up = i32::try_from(at).expect("a register save area under two gigabytes");
2994 let mem = mir::Mem::at(mir::Operand::read(base, self.gpr)).plus(up);
2995 self.out.build(out, store).uses(reg, class).mem(mem).finish();
2996 }
2997 }
2998
2999 /// The address of one of the function's stack objects, in a fresh register.
3000 ///
3001 /// Written with nothing in its displacement, because where an object is in a frame is not known
3002 /// until after allocation, and given to [`crate::finish`] to fill in the way an `alloca` is.
3003 fn frame_address(&mut self, out: mir::Block, local: usize) -> mir::Reg {
3004 let reg = self.out.new_vreg(self.gpr);
3005 let lea = mir::Opcode::new(self.names.intern(&format!("{PREFIX}{}", x86_64::FRAME.lea)));
3006 let sp = mir::Operand::read(mir::Reg::physical(self.conv.stack_pointer), self.gpr);
3007 let made = self.out.build(out, lea).def(reg, self.gpr).mem(mir::Mem::at(sp)).finish();
3008 self.stack.addresses.push((made, local));
3009 reg
3010 }
3011
3012 /// Whether an instruction is one no machine instruction is written for where it stands.
3013 ///
3014 /// Four of them, and none is a lowering decision, which is why none is a rule. A constant is
3015 /// written where a register for it is first wanted rather than where the IR put it, and every
3016 /// reader of one may have folded it into an immediate, in which case nowhere is the right
3017 /// place. A return of nothing has nothing to put anywhere: the epilogue gives the frame back
3018 /// and leaves, and it is appended to every block with no successors long after this has
3019 /// finished, so a return with a value is one instruction here and a return without one is
3020 /// none. Unless the value went back through memory, in which case there is something to put
3021 /// somewhere after all and the IR does not carry it: the address the caller handed over has
3022 /// to be in `rax` on the way out, and [`Lowering::returned`] is what writes that.
3023 ///
3024 /// An unconditional jump is the third, and there is even less of it: the edge is on the
3025 /// block, and whether the block it goes to is the next one and needs no jump at all is the
3026 /// block layout's answer rather than this one's.
3027 ///
3028 /// The fourth is a point control does not arrive at, in both of the forms the IR has for it:
3029 /// the `unreachable` terminator the front end puts at the end of a function whose body can run
3030 /// off the bottom, and the `unreachable_hint` a call to `__builtin_unreachable` becomes. What
3031 /// to write for a place nothing reaches is a question with no wrong answer, and nothing is the
3032 /// smallest one and the one gcc 16.2.0 gives at `-O0`. The terminator leaves the block with no
3033 /// successors, so the epilogue lands at the end of it the way it does on any other block that
3034 /// goes nowhere, and the function cannot fall out of its own last instruction into whatever
3035 /// the assembler puts next.
3036 fn writes_nothing(&self, inst: Inst) -> bool {
3037 let data = &self.source[inst];
3038 match data.opcode {
3039 Opcode::IConst | Opcode::Jump | Opcode::Unreachable | Opcode::UnreachableHint => true,
3040 Opcode::Return => self.source[data.args].is_empty() && self.sret().is_none(),
3041 _ => false,
3042 }
3043 }
3044
3045 /// What every instruction in one block matched, with a set of values nobody may take.
3046 ///
3047 /// Backwards, because an instruction that has been folded into a later one does not get to
3048 /// fold anything into itself: the rule that took it only reached one level down, so what is
3049 /// under it is not in the term the matcher saw and cannot be replaced.
3050 fn decide(&self, insts: &[Inst], refused: &HashSet<Value>) -> Decided {
3051 let mut found: Vec<Option<Match<Term>>> = (0..insts.len()).map(|_| None).collect();
3052 let mut plans: Vec<Option<Plan>> = vec![None; insts.len()];
3053 let mut folded: Vec<Inst> = Vec::new();
3054 for (index, &inst) in insts.iter().enumerate().rev() {
3055 if folded.contains(&inst) {
3056 continue;
3057 }
3058 if let Some((plan, matched)) = self.select(inst, refused) {
3059 folded.extend(self.folds(inst, plan));
3060 found[index] = Some(matched);
3061 plans[index] = Some(plan);
3062 }
3063 }
3064 Decided { found, plans, folded }
3065 }
3066
3067 /// A value some of its readers took and some of them did not, which is the one case folding
3068 /// buys nothing.
3069 ///
3070 /// Folding does not delete the instruction that computed a value for anybody else, so a
3071 /// reader that did not take it still needs it in a register and the instruction stays. The
3072 /// reader that did take it now does that work again. Either all of them take it, in which
3073 /// case nothing is left to read it and the instruction goes, or none of them do.
3074 ///
3075 /// The count is over the whole function rather than over the block, since a value read from
3076 /// another block is read from a register there whatever this block decides. An instruction
3077 /// built by name rather than matched, a call being the one that matters, has no plan and so
3078 /// takes nothing, which is the right answer for it as well.
3079 fn left_alive(&self, insts: &[Inst], plans: &[Option<Plan>]) -> Option<Value> {
3080 let mut taken = vec![0u32; self.uses.len()];
3081 for (&inst, plan) in insts.iter().zip(plans) {
3082 let Some(plan) = plan else { continue };
3083 let args = &self.source[self.source[inst].args];
3084 for (index, &arg) in args.iter().take(MAX_ARGS).enumerate() {
3085 if plan[index] == Shown::Expand {
3086 taken[arg.index()] += 1;
3087 }
3088 }
3089 }
3090 for (&inst, plan) in insts.iter().zip(plans) {
3091 let Some(plan) = plan else { continue };
3092 let args = &self.source[self.source[inst].args];
3093 for (index, &arg) in args.iter().take(MAX_ARGS).enumerate() {
3094 if plan[index] == Shown::Expand && taken[arg.index()] < self.uses[arg.index()] {
3095 return Some(arg);
3096 }
3097 }
3098 }
3099 None
3100 }
3101
3102 /// The rule that fires on an instruction, and what it bound.
3103 ///
3104 /// The plans are tried in order and the first that matches wins, which is the maximal munch
3105 /// `spec/10-backend.md` asks for: a plan that offers more to the matcher is tried before one
3106 /// that offers less.
3107 fn select(&self, inst: Inst, refused: &HashSet<Value>) -> Option<(Plan, Match<Term>)> {
3108 for plan in self.plans(inst, refused) {
3109 let terms = Terms::new(self.source, inst, plan);
3110 if let Some(matched) = TABLE.find(&terms, Term::Root) {
3111 return Some((plan, matched));
3112 }
3113 }
3114 None
3115 }
3116
3117 /// Every way this instruction can be shown to the matcher, most offered first.
3118 fn plans(&self, inst: Inst, refused: &HashSet<Value>) -> Vec<Plan> {
3119 let args = &self.source[self.source[inst].args];
3120 let mut plans = vec![PLAIN];
3121 for (index, &arg) in args.iter().enumerate().take(MAX_ARGS) {
3122 let mut ways = Vec::new();
3123 if self.foldable(inst, arg, refused) {
3124 ways.push(Shown::Expand);
3125 }
3126 if Terms::new(self.source, inst, PLAIN).constant(arg).is_some() {
3127 ways.push(Shown::Const);
3128 }
3129 ways.push(Shown::Reg);
3130 plans = plans
3131 .into_iter()
3132 .flat_map(|plan| {
3133 ways.iter().map(move |&way| {
3134 let mut next = plan;
3135 next[index] = way;
3136 next
3137 })
3138 })
3139 .collect();
3140 }
3141 plans
3142 }
3143
3144 /// Whether an operand may be shown as the instruction that computed it.
3145 ///
3146 /// It has to be in the same block, because a rule that folds one instruction into another
3147 /// moves the work to where the second one is. It has to be something rather than a block
3148 /// parameter, and not a constant, which is shown as a constant instead. And it has to be a
3149 /// value [`Lowering::left_alive`] has not put back, which is how the one reader at a time
3150 /// question is asked here: this says yes to a value with any number of readers, and a value
3151 /// only some of them could take is refused after the fact and asked again.
3152 ///
3153 /// A value with several readers used to be refused outright, on the reasoning that folding
3154 /// does not delete the instruction for anybody else. That reasoning is about the set of
3155 /// readers and was being applied to one reader at a time, which is stricter than it needs to
3156 /// be: when every reader takes it there is nobody left to read it and the instruction goes.
3157 /// An address a store and a load share is the shape that matters, since a memory operand has
3158 /// room for the whole of it and both readers have a memory operand.
3159 fn foldable(&self, into: Inst, value: Value, refused: &HashSet<Value>) -> bool {
3160 let Def::Result { inst, .. } = self.source[value].def else { return false };
3161 if self.source[inst].opcode == Opcode::IConst || refused.contains(&value) {
3162 return false;
3163 }
3164 self.source.block_of(inst).is_some()
3165 && self.source.block_of(inst) == self.source.block_of(into)
3166 }
3167
3168 /// The instructions a match folded into the one it matched.
3169 ///
3170 /// The plan is what says this, not the bindings: a binding is a register or a number either
3171 /// way, and an operand shown as the instruction that computed it is one no rule could have
3172 /// matched without taking that instruction, because the plan offered the matcher nothing
3173 /// else to call it.
3174 fn folds(&self, inst: Inst, plan: Plan) -> Vec<Inst> {
3175 let args = &self.source[self.source[inst].args];
3176 args.iter()
3177 .take(MAX_ARGS)
3178 .enumerate()
3179 .filter(|&(index, _)| plan[index] == Shown::Expand)
3180 .filter_map(|(_, &arg)| match self.source[arg].def {
3181 Def::Result { inst, .. } => Some(inst),
3182 Def::Param { .. } => None,
3183 })
3184 .collect()
3185 }
3186
3187 /// Build the machine instruction a match calls for.
3188 fn emit(&mut self, inst: Inst, matched: &Match<Term>) -> Result<(), Unsupported> {
3189 let rule: &Rule = TABLE.rule(matched);
3190 let pieces = rule.replacement;
3191 let Some(Piece::App { head, arity }) = pieces.first() else {
3192 return Err(self.unsupported(inst));
3193 };
3194 let opcode = head.strip_prefix(PREFIX).ok_or_else(|| self.unsupported(inst))?;
3195 let form = x86_64::form(opcode).ok_or_else(|| self.unsupported(inst))?;
3196
3197 let mut read = Read::default();
3198 let mut at = 1;
3199 for _ in 0..*arity {
3200 at = self.read(inst, pieces, at, &matched.bindings, &mut read)?;
3201 }
3202
3203 let descs = form.operands();
3204 let writes = descs.iter().take_while(|desc| desc.role.is_def()).count();
3205 if descs.len() - writes != read.regs.len() {
3206 return Err(self.unsupported(inst));
3207 }
3208
3209 // The first thing the instruction writes is what it computes, and any others are
3210 // registers the machine destroys on the way, which are fresh because nothing else is in
3211 // them and nothing reads them. An instruction that writes nothing at all is one whose
3212 // whole purpose is its effect, which is what a store is, and there is no result to put
3213 // anywhere.
3214 let mut regs = Vec::new();
3215 if writes > 0 {
3216 let result = self.source[inst].first_result.ok_or_else(|| self.unsupported(inst))?;
3217 regs.push(self.new_reg(result));
3218 // The rest are the registers the machine destroys on the way, and the class each is in
3219 // is the one the instruction's description gives it rather than a guess, so that an
3220 // instruction that wrecks a register in the other file says so.
3221 regs.extend(descs[1..writes].iter().map(|desc| self.out.new_vreg(desc.class)));
3222 } else if self.source[inst].first_result.is_some() {
3223 // A rule that throws away a value the IR gave a name to would leave every reader of
3224 // that name with nothing to read, so it is a rule this and the target disagree about.
3225 return Err(self.unsupported(inst));
3226 }
3227 regs.extend(read.regs.iter().copied());
3228
3229 let block = self.at.expect("a block is being filled");
3230 let opcode = mir::Opcode::new(self.names.intern(head));
3231 let mut build = self.out.build(block, opcode).at(self.source.span(inst));
3232 for (desc, reg) in descs.iter().zip(regs) {
3233 let operand = mir::Operand {
3234 reg,
3235 class: desc.class,
3236 role: desc.role,
3237 constraint: desc.constraint,
3238 };
3239 build = build.operand(operand);
3240 }
3241 if let Some(mem) = read.mem {
3242 build = build.mem(mem);
3243 }
3244 if let Some(imm) = read.imm {
3245 build = build.imm(imm);
3246 }
3247 build.finish();
3248 Ok(())
3249 }
3250
3251 /// Read one argument of a replacement, which is a register, a number or an address.
3252 ///
3253 /// Gives back the position after it, because a replacement is flat and an address takes
3254 /// arguments of its own.
3255 fn read(
3256 &mut self,
3257 inst: Inst,
3258 pieces: &'static [Piece],
3259 at: usize,
3260 bindings: &[Term],
3261 out: &mut Read,
3262 ) -> Result<usize, Unsupported> {
3263 match pieces.get(at) {
3264 Some(Piece::Int(value)) => {
3265 out.imm = i64::try_from(*value).ok();
3266 Ok(at + 1)
3267 }
3268 Some(Piece::Var { index, .. }) => {
3269 match bindings.get(*index) {
3270 Some(&Term::Reg(value)) => {
3271 let reg = self.reg_of(value)?;
3272 out.regs.push(reg);
3273 }
3274 Some(&Term::Num(value)) => out.imm = i64::try_from(value).ok(),
3275 // A pattern binds a register or a number and nothing else, so this is a
3276 // rule the matcher and this file disagree about.
3277 _ => return Err(self.unsupported(inst)),
3278 }
3279 Ok(at + 1)
3280 }
3281 Some(Piece::App { head, arity }) => {
3282 let kind = x86_64::address(head).ok_or_else(|| self.unsupported(inst))?;
3283 let mut inner = Read::default();
3284 let mut next = at + 1;
3285 for _ in 0..*arity {
3286 next = self.read(inst, pieces, next, bindings, &mut inner)?;
3287 }
3288 let mem = address(kind, &inner, self.gpr).ok_or_else(|| self.unsupported(inst))?;
3289 out.mem = Some(mem);
3290 Ok(next)
3291 }
3292 None => Err(self.unsupported(inst)),
3293 }
3294 }
3295
3296 /// The register a value is in, materializing it if it is a constant that has not been put in
3297 /// one yet.
3298 ///
3299 /// A constant is written where it is wanted rather than where the IR defined it, and where it
3300 /// is wanted is a block that need not be the one the IR defined it in. So the register holding
3301 /// one is only good inside the block it was written into, and a second block that wants the
3302 /// same constant gets its own. Anything else is a register read where nothing wrote it: the
3303 /// IR guarantees a definition dominates its uses, and this moved the definition.
3304 ///
3305 /// Writing the number again is also the right answer and not merely the safe one. It is one
3306 /// instruction that reads nothing, which is cheaper than holding a register live across a
3307 /// branch for it, and it is what a rematerializing allocator would do with the value anyway.
3308 fn reg_of(&mut self, value: Value) -> Result<mir::Reg, Unsupported> {
3309 let constant = match self.source[value].def {
3310 Def::Result { inst, .. } => {
3311 (self.source[inst].opcode == Opcode::IConst).then_some(inst)
3312 }
3313 Def::Param { .. } => None,
3314 };
3315 let here = self.at.expect("a block is being filled");
3316 if let Some(reg) = self.regs[value.index()] {
3317 if constant.is_none() || self.written[value.index()] == Some(here) {
3318 return Ok(reg);
3319 }
3320 }
3321 if let Some(inst) = constant {
3322 // Cleared so that the register the constant is written into is a new one rather than
3323 // the one the block above wrote, which is still being read up there.
3324 self.regs[value.index()] = None;
3325 // Nothing is refused here. A constant is written on its own, out of the loop over the
3326 // block, and the operands of the rule that writes one are the number and nothing else.
3327 let matched = self
3328 .select(inst, &HashSet::new())
3329 .map(|(_, matched)| matched)
3330 .ok_or_else(|| self.unsupported(inst))?;
3331 self.emit(inst, &matched)?;
3332 // The same mark the loop over the instructions makes, and it has to be made here as
3333 // well because this is the only place a constant is ever selected: the loop skips one
3334 // where the IR wrote it, so a rule that lowers a constant fires from nowhere else and
3335 // would be reported as a rule nothing reaches.
3336 self.fired.mark(matched.rule);
3337 self.written[value.index()] = Some(here);
3338 return Ok(self.regs[value.index()].expect("a constant is written into a register"));
3339 }
3340 Ok(self.new_reg(value))
3341 }
3342
3343 /// Which register file a value of that type lives in.
3344 ///
3345 /// The vector one for the two float widths the machine has scalar instructions for and for the
3346 /// one it only moves, and the general purpose one for everything else. An eighty bit `long
3347 /// double` is in neither, and it is here rather than in the vector class on purpose: it would
3348 /// be put in a register that cannot hold it, and there is no rule that names one, so the
3349 /// instruction computing it is reported. The wrong class would make that a wrong program
3350 /// instead of a refused one.
3351 ///
3352 /// A hundred and twenty eight bit float is in the vector class and fits it exactly, which is
3353 /// the difference. Nothing computes in it, so every arithmetic on one is still reported, and
3354 /// what the class buys is the moves: a register that holds the whole value is a register a
3355 /// spill, a reload and a copy are each one instruction for.
3356 fn class_of(&self, ty: Type) -> RegClass {
3357 if crate::term::in_vector_file(ty) { self.conv.sse_class } else { self.gpr }
3358 }
3359
3360 /// A fresh register for a value, which is what the instruction computing it writes.
3361 fn new_reg(&mut self, value: Value) -> mir::Reg {
3362 if let Some(reg) = self.regs[value.index()] {
3363 return reg;
3364 }
3365 let reg = self.out.new_vreg(self.class_of(self.source[value].ty));
3366 self.regs[value.index()] = Some(reg);
3367 reg
3368 }
3369
3370 fn unsupported(&self, inst: Inst) -> Unsupported {
3371 let data = &self.source[inst];
3372 Unsupported::Inst {
3373 inst,
3374 term: Terms::new(self.source, inst, PLAIN).name(inst),
3375 opcode: data.opcode,
3376 ty: data.first_result.map(|result| self.source[result].ty),
3377 }
3378 }
3379}
3380
3381/// What the arguments of one replacement came to.
3382#[derive(Debug, Default)]
3383struct Read {
3384 regs: Vec<mir::Reg>,
3385 imm: Option<i64>,
3386 mem: Option<mir::Mem>,
3387}
3388
3389/// The addressing mode an address constructor's arguments make.
3390///
3391/// One arm per constructor rather than a question asked of the kind, because what the arguments
3392/// mean is the whole of what tells the four apart: the same register is a base in one and an
3393/// index in another, and the same constant is a scale in one and a displacement in another.
3394fn address(kind: x86_64::Address, read: &Read, gpr: RegClass) -> Option<mir::Mem> {
3395 let mut regs = read.regs.iter().copied().map(|reg| mir::Operand::read(reg, gpr));
3396 match kind {
3397 x86_64::Address::BaseIndexScale => {
3398 let base = regs.next()?;
3399 let index = regs.next()?;
3400 Some(mir::Mem::at(base).indexed(index, u8::try_from(read.imm?).ok()?))
3401 }
3402 x86_64::Address::IndexScale => Some(mir::Mem {
3403 base: None,
3404 index: Some(regs.next()?),
3405 scale: u8::try_from(read.imm?).ok()?,
3406 disp: 0,
3407 symbol: None,
3408 block: None,
3409 reach: mir::Reach::Itself,
3410 segment: None,
3411 }),
3412 x86_64::Address::Base => Some(mir::Mem::at(regs.next()?)),
3413 // The rule that writes this has a guard saying the constant fits, so a displacement that
3414 // does not is a rule and a target that disagree rather than a program this cannot compile.
3415 x86_64::Address::BaseOffset => {
3416 Some(mir::Mem { disp: i32::try_from(read.imm?).ok()?, ..mir::Mem::at(regs.next()?) })
3417 }
3418 }
3419}
3420
3421/// The table this selector matches with.
3422///
3423/// One target for now, because one target has a rule file. Which table to use becomes a question
3424/// the moment a second one does, and the answer will be the target the session was given rather
3425/// than a constant here.
3426static TABLE: &Table = &crate::select::x86_64::TABLE;
3427
3428#[cfg(test)]
3429mod tests {
3430 use rucc_ir::{
3431 AsmInfo, Builder, CallInfo, Flags, InstData, MemInfo, MemOrder, Restrict, Signature, Type,
3432 };
3433 use rucc_regalloc::assign::Env;
3434 use rucc_target::x86_64::{FRAME, REGS, SYSV};
3435
3436 use super::*;
3437 use crate::finish::{Convention, finish};
3438 use crate::frame::{Frame, Incoming, Layout};
3439
3440 /// A function of as many 64 bit parameters as the test wants, and the block they are in.
3441 fn blank(params: &[Type]) -> (Interner, Func, Block, Vec<Value>) {
3442 let mut names = Interner::new();
3443 let mut func = Func::new(names.intern("f"), Signature::new());
3444 let block = func.create_block();
3445 let values = params.iter().map(|&ty| func.append_param(block, ty)).collect();
3446 (names, func, block, values)
3447 }
3448
3449 /// An ordinary access: not atomic, and aligned enough that nothing here has an opinion.
3450 /// Neither field reaches selection, which is the point of saying it once here.
3451 fn plain() -> MemInfo {
3452 MemInfo {
3453 size: 0,
3454 align: 1,
3455 order: MemOrder::NotAtomic,
3456 tbaa: None,
3457 owns: 0,
3458 restrict: Restrict::NONE,
3459 }
3460 }
3461
3462 /// What the allocator is given: every integer register the convention offers except two, held
3463 /// back so that a move on an edge has somewhere to break a cycle and a spilled value has
3464 /// somewhere to be read into. Which two does not matter, and holding back the last two the
3465 /// convention would reach for leaves every expectation below unchanged.
3466 fn env() -> Env {
3467 const SCRATCH: [PhysReg; 2] = [x86_64::R10, x86_64::R11];
3468 let order: Vec<PhysReg> =
3469 SYSV.int_order.iter().copied().filter(|reg| !SCRATCH.contains(reg)).collect();
3470 Env::new().with(x86_64::GPR, &order, &SCRATCH)
3471 }
3472
3473 /// The machine IR text a function lowers to.
3474 fn lower(names: &mut Interner, source: &Func) -> String {
3475 let out = func(source, names, &SYSV, &Elsewhere::default())
3476 .expect("every instruction has a rule");
3477 mir::print_func(&out.func, names, ®S)
3478 }
3479
3480 #[test]
3481 fn an_addition_of_two_registers_is_one_instruction() {
3482 let i32 = Type::int(32);
3483 let (mut names, mut func, block, args) = blank(&[i32, i32]);
3484 let mut build = Builder::new(&mut func, block);
3485 build.binary(Opcode::Add, args[0], args[1], Flags::default());
3486
3487 assert_eq!(
3488 lower(&mut names, &func),
3489 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
3490 %1:gpr($rsi) = x64.arg_val_32\n %2:gpr(reuse 1) = x64.add_rr_32 %0, %1\n}\n"
3491 );
3492 }
3493
3494 #[test]
3495 fn a_constant_operand_becomes_an_immediate() {
3496 let i32 = Type::int(32);
3497 let (mut names, mut func, block, args) = blank(&[i32]);
3498 let mut build = Builder::new(&mut func, block);
3499 let seven = build.iconst(i32, 7);
3500 build.binary(Opcode::Add, args[0], seven, Flags::default());
3501
3502 // The constant is in the instruction and nothing was written to hold it, which is what
3503 // materializing one where a register for it is wanted buys.
3504 assert_eq!(
3505 lower(&mut names, &func),
3506 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
3507 %1:gpr(reuse 1) = x64.add_ri_32 %0, 7\n}\n"
3508 );
3509 }
3510
3511 #[test]
3512 fn a_constant_too_wide_for_an_immediate_goes_into_a_register() {
3513 let i64 = Type::int(64);
3514 let (mut names, mut func, block, args) = blank(&[i64]);
3515 let mut build = Builder::new(&mut func, block);
3516 let big = build.iconst(i64, i128::from(i32::MAX) + 1);
3517 build.binary(Opcode::Add, args[0], big, Flags::default());
3518
3519 // Nobody wrote this fallback down. The rule that takes an immediate has a guard that
3520 // turns a number this wide down, so it does not fire, and the next way of showing the
3521 // operand puts it in a register.
3522 assert_eq!(
3523 lower(&mut names, &func),
3524 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
3525 %1:gpr = x64.mov_ri_64 2147483648\n %2:gpr(reuse 1) = x64.add_rr_64 %0, %1\n}\n"
3526 );
3527 }
3528
3529 #[test]
3530 fn an_index_calculation_folds_into_an_address() {
3531 let i64 = Type::int(64);
3532 let (mut names, mut func, block, args) = blank(&[i64, i64]);
3533 let mut build = Builder::new(&mut func, block);
3534 let four = build.iconst(i64, 4);
3535 let scaled = build.binary(Opcode::Mul, args[1], four, Flags::default());
3536 build.binary(Opcode::Add, args[0], scaled, Flags::default());
3537
3538 // Three IR instructions and one machine instruction. The multiply is gone because the
3539 // rule that matched reached down and took it.
3540 assert_eq!(
3541 lower(&mut names, &func),
3542 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
3543 %1:gpr($rsi) = x64.arg_val_64\n %2:gpr = x64.lea_64 [%0 + %1*4]\n}\n"
3544 );
3545 }
3546
3547 #[test]
3548 fn an_instruction_every_reader_can_take_is_folded_into_all_of_them() {
3549 let i64 = Type::int(64);
3550 let (mut names, mut func, block, args) = blank(&[i64, i64]);
3551 let mut build = Builder::new(&mut func, block);
3552 let four = build.iconst(i64, 4);
3553 let scaled = build.binary(Opcode::Mul, args[1], four, Flags::default());
3554 let first = build.binary(Opcode::Add, args[0], scaled, Flags::default());
3555 build.binary(Opcode::Add, first, scaled, Flags::default());
3556
3557 // Both readers have room for a scaled index, so both of them take it and nothing is left
3558 // to read the multiply. Three IR instructions become two machine ones, where refusing to
3559 // fold into either reader would have left three.
3560 assert_eq!(
3561 lower(&mut names, &func),
3562 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
3563 %1:gpr($rsi) = x64.arg_val_64\n %2:gpr = x64.lea_64 [%0 + %1*4]\n \
3564 %3:gpr = x64.lea_64 [%2 + %1*4]\n}\n"
3565 );
3566 }
3567
3568 #[test]
3569 fn an_instruction_one_of_its_readers_cannot_take_is_folded_into_none_of_them() {
3570 let i64 = Type::int(64);
3571 let (mut names, mut func, block, args) = blank(&[i64, i64]);
3572 let mut build = Builder::new(&mut func, block);
3573 let four = build.iconst(i64, 4);
3574 let scaled = build.binary(Opcode::Mul, args[1], four, Flags::default());
3575 build.binary(Opcode::Add, args[0], scaled, Flags::default());
3576 build.store(scaled, args[0], plain(), Flags::default());
3577
3578 // The addition has room for the multiply and the store does not: what a store writes is
3579 // a register, and no rule reaches through it. Folding into the addition alone would
3580 // leave the multiply where it is for the store to read and do the work twice, so the
3581 // multiply is put back and both readers read the register it wrote.
3582 let text = lower(&mut names, &func);
3583 assert!(text.contains("x64.lea_64 [%1*4]"), "{text}");
3584 assert!(text.contains("x64.add_rr_64"), "{text}");
3585 }
3586
3587 #[test]
3588 fn a_shift_by_a_register_asks_for_it_in_cl() {
3589 let i32 = Type::int(32);
3590 let (mut names, mut func, block, args) = blank(&[i32, i32]);
3591 let mut build = Builder::new(&mut func, block);
3592 build.binary(Opcode::Shl, args[0], args[1], Flags::default());
3593
3594 // The fixed register is not in the rule. It is what the target says the instruction does
3595 // with its operands, and the allocator is what will act on it.
3596 let text = lower(&mut names, &func);
3597 assert!(text.contains("x64.shl_rcl_32 %0, %1($rcx)"), "{text}");
3598 }
3599
3600 #[test]
3601 fn a_division_names_the_registers_and_the_register_it_destroys() {
3602 let i32 = Type::int(32);
3603 let (mut names, mut func, block, args) = blank(&[i32, i32]);
3604 let mut build = Builder::new(&mut func, block);
3605 build.binary(Opcode::SDiv, args[0], args[1], Flags::default());
3606
3607 // Two definitions, because a division writes the remainder whether anybody wanted it or
3608 // not, and the second one is early because it is destroyed before the operands are read.
3609 let text = lower(&mut names, &func);
3610 assert!(
3611 text.contains("%2:gpr($rax), early %3:gpr($rdx) = x64.idiv_quo_32 %0($rax), %1"),
3612 "{text}"
3613 );
3614 }
3615
3616 #[test]
3617 fn a_load_reads_through_the_register_the_address_is_in() {
3618 let i64 = Type::int(64);
3619 let (mut names, mut func, block, args) = blank(&[i64]);
3620 let mut build = Builder::new(&mut func, block);
3621 build.load(Type::int(32), args[0], plain(), Flags::default());
3622
3623 assert_eq!(
3624 lower(&mut names, &func),
3625 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
3626 %1:gpr = x64.mov_rm_32 [%0]\n}\n"
3627 );
3628 }
3629
3630 #[test]
3631 fn a_store_writes_no_register_and_the_value_it_writes_is_the_one_the_ir_gave_it() {
3632 let (mut names, mut func, block, args) = blank(&[Type::int(32), Type::int(64)]);
3633 let mut build = Builder::new(&mut func, block);
3634 build.store(args[0], args[1], plain(), Flags::default());
3635
3636 // The value is the first parameter and the address is the second, and the instruction
3637 // takes them the other way round. Getting that backwards would compile to a store of the
3638 // address into the value, which is a program that runs and does the wrong thing.
3639 assert_eq!(
3640 lower(&mut names, &func),
3641 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
3642 %1:gpr($rsi) = x64.arg_val_64\n x64.mov_mr_32 %0, [%1]\n}\n"
3643 );
3644 }
3645
3646 #[test]
3647 fn an_address_with_a_constant_added_folds_into_the_access() {
3648 let i64 = Type::int(64);
3649 let (mut names, mut func, block, args) = blank(&[i64]);
3650 let mut build = Builder::new(&mut func, block);
3651 let twelve = build.iconst(i64, 12);
3652 let field = build.binary(Opcode::Add, args[0], twelve, Flags::default());
3653 build.load(Type::int(64), field, plain(), Flags::default());
3654
3655 // Two IR instructions and one machine instruction, which is what every read of a field
3656 // of a structure comes to.
3657 assert_eq!(
3658 lower(&mut names, &func),
3659 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
3660 %1:gpr = x64.mov_rm_64 [%0 + 12]\n}\n"
3661 );
3662 }
3663
3664 #[test]
3665 fn a_displacement_too_wide_to_encode_leaves_the_addition_where_it_is() {
3666 let i64 = Type::int(64);
3667 let (mut names, mut func, block, args) = blank(&[i64]);
3668 let mut build = Builder::new(&mut func, block);
3669 let big = build.iconst(i64, i128::from(i32::MAX) + 1);
3670 let far = build.binary(Opcode::Add, args[0], big, Flags::default());
3671 build.load(Type::int(32), far, plain(), Flags::default());
3672
3673 // A displacement is signed and 32 bits. The rule that folds one has a guard that turns
3674 // this down, so the addition stays and the load reads through what it produced. Nobody
3675 // wrote that fallback: it is the next way of showing the operand.
3676 let text = lower(&mut names, &func);
3677 assert!(text.contains("x64.mov_rm_32 [%2]"), "{text}");
3678 assert!(text.contains("x64.add_rr_64"), "{text}");
3679 }
3680
3681 #[test]
3682 fn a_store_of_a_value_that_was_loaded_is_two_instructions_and_no_arithmetic() {
3683 let i64 = Type::int(64);
3684 let (mut names, mut func, block, args) = blank(&[i64, i64]);
3685 let mut build = Builder::new(&mut func, block);
3686 let got = build.load(Type::int(8), args[0], plain(), Flags::default());
3687 build.store(got, args[1], plain(), Flags::default());
3688
3689 // A load feeding a store is the one place folding would be wrong: an x86-64 `mov` has at
3690 // most one memory operand, and there is no rule that takes two, so the load is left where
3691 // it is and the store reads the register it wrote.
3692 assert_eq!(
3693 lower(&mut names, &func),
3694 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
3695 %1:gpr($rsi) = x64.arg_val_64\n %2:gpr = x64.mov_rm_8 [%0]\n \
3696 x64.mov_mr_8 %2, [%1]\n}\n"
3697 );
3698 }
3699
3700 #[test]
3701 fn an_access_at_a_width_no_rule_is_written_at_is_reported() {
3702 let i64 = Type::int(64);
3703 let (mut names, mut source, block, args) = blank(&[i64]);
3704 let mut build = Builder::new(&mut source, block);
3705 build.load(Type::int(128), args[0], plain(), Flags::default());
3706
3707 // The width is the whole of what is wrong here, so the width is in the message: `load`
3708 // on its own is written about at every other width and would send a reader looking in
3709 // the wrong place.
3710 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
3711 .expect_err("nothing loads 128 bits");
3712 assert_eq!(failed.to_string(), "no rule lowers a `load` producing a `i128`");
3713 }
3714
3715 #[test]
3716 fn a_return_asks_for_the_value_in_the_register_the_caller_reads() {
3717 let (mut names, mut func, block, args) = blank(&[Type::int(32)]);
3718 let mut build = Builder::new(&mut func, block);
3719 build.ret(&[args[0]]);
3720
3721 // The register is not in the rule, the same way `cl` is not in the rule for a shift. It
3722 // is what the target says the instruction does with its operand, and the allocator is
3723 // what will act on it. There is no `ret` here, because giving the frame back has to
3724 // happen between this and leaving and the frame is not worked out yet.
3725 assert_eq!(
3726 lower(&mut names, &func),
3727 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
3728 x64.ret_val_32 %0($rax)\n}\n"
3729 );
3730 }
3731
3732 #[test]
3733 fn a_return_of_two_values_asks_for_the_second_register_as_well() {
3734 let i64 = Type::int(64);
3735 let (mut names, mut func, block, args) = blank(&[i64, i64]);
3736 let mut build = Builder::new(&mut func, block);
3737 build.ret(&[args[0], args[1]]);
3738
3739 // `struct { long a, b; } f(long a, long b)`, after the front end has classified it. Both
3740 // halves are integers, so the second is in the second integer return register, and both
3741 // pseudos say so the same way the one for a single value does.
3742 assert_eq!(
3743 lower(&mut names, &func),
3744 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
3745 %1:gpr($rsi) = x64.arg_val_64\n x64.ret_val_64 %0($rax)\n \
3746 x64.ret_val2_64 %1($rdx)\n}\n"
3747 );
3748 }
3749
3750 #[test]
3751 fn two_values_back_in_different_files_are_both_the_first_of_their_own() {
3752 let f64 = Type::float(rucc_ir::Float::F64);
3753 let (mut names, mut func, block, args) = blank(&[f64, Type::int(64)]);
3754 let mut build = Builder::new(&mut func, block);
3755 build.ret(&[args[0], args[1]]);
3756
3757 // `struct { double a; long b; } f(double a, long b)`. The two files are counted apart, so
3758 // neither half is the second of anything and the `double` is in `xmm0` rather than in the
3759 // register a second `double` would have been in. Getting this wrong is not a crash: the
3760 // caller reads a register nobody wrote, and this is where that is ruled out.
3761 assert_eq!(
3762 lower(&mut names, &func),
3763 "mfunc @f {\nblock0:\n %0:xmm($xmm0) = x64.arg_val_f64\n \
3764 %1:gpr($rdi) = x64.arg_val_64\n x64.ret_val_f64 %0($xmm0)\n \
3765 x64.ret_val_64 %1($rax)\n}\n"
3766 );
3767 }
3768
3769 #[test]
3770 fn two_of_the_same_file_back_take_the_first_two_of_it() {
3771 let f64 = Type::float(rucc_ir::Float::F64);
3772 let (mut names, mut func, block, args) = blank(&[f64, f64]);
3773 let mut build = Builder::new(&mut func, block);
3774 build.ret(&[args[0], args[1]]);
3775
3776 // `struct { double x, y; } f(double x, double y)`, which is the vector half of the pair
3777 // above and counts in its own file the same way.
3778 assert_eq!(
3779 lower(&mut names, &func),
3780 "mfunc @f {\nblock0:\n %0:xmm($xmm0) = x64.arg_val_f64\n \
3781 %1:xmm($xmm1) = x64.arg_val_f64\n x64.ret_val_f64 %0($xmm0)\n \
3782 x64.ret_val2_f64 %1($xmm1)\n}\n"
3783 );
3784 }
3785
3786 /// A function whose answer goes back through memory, with the pointer to the space for it in
3787 /// front of whatever else it takes. Only the signature says it is one.
3788 fn returning_through_memory(params: &[Type]) -> (Interner, Func, Block, Vec<Value>) {
3789 let mut names = Interner::new();
3790 let sret = Abi::Sret { size: 32, align: 8 };
3791 let mut signature = Signature::new().and_param(Param::with_abi(Type::PTR, sret));
3792 signature.params.extend(params.iter().copied().map(Param::new));
3793 let mut func = Func::new(names.intern("f"), signature);
3794 let block = func.create_block();
3795 let space = func.append_param(block, Type::PTR);
3796 let values = std::iter::once(space)
3797 .chain(params.iter().map(|&ty| func.append_param(block, ty)))
3798 .collect();
3799 (names, func, block, values)
3800 }
3801
3802 #[test]
3803 fn the_space_a_return_through_memory_was_given_goes_back_in_the_first_return_register() {
3804 let (mut names, mut func, block, _) = returning_through_memory(&[]);
3805 Builder::new(&mut func, block).ret(&[]);
3806
3807 // `struct big f(void)`, where `big` is too large to come back in registers. The `return`
3808 // carries nothing, because the value went into the space the caller handed over, and the
3809 // document still says that address comes back in `rax`. Nothing in the IR says it, so the
3810 // convention says it, and the pseudo is the one any other pointer return would use.
3811 assert_eq!(
3812 lower(&mut names, &func),
3813 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
3814 x64.ret_val_64 %0($rax)\n}\n"
3815 );
3816 }
3817
3818 #[test]
3819 fn what_the_function_did_in_between_does_not_take_the_register_off_it() {
3820 let (mut names, mut func, block, args) = returning_through_memory(&[Type::int(32)]);
3821 let mut build = Builder::new(&mut func, block);
3822 build.store(args[1], args[0], plain(), Flags::default());
3823 build.ret(&[]);
3824
3825 // The register is a read at the end and not a move at the start, so it is live across
3826 // everything between the two and the allocator has to keep it somewhere. In a function
3827 // with a call in it that somewhere is a callee saved register, and the address comes back
3828 // into `rax` here rather than whatever the last instruction happened to leave there. That
3829 // is issue #333, and a store is enough to show the value outlives the entry block.
3830 let text = lower(&mut names, &func);
3831 assert!(text.contains("x64.mov_mr_32 %1, [%0]"), "{text}");
3832 assert!(text.ends_with(" x64.ret_val_64 %0($rax)\n}\n"), "{text}");
3833 }
3834
3835 #[test]
3836 fn a_pointer_that_is_only_a_pointer_is_not_given_back() {
3837 let (mut names, mut func, block, args) = blank(&[Type::PTR]);
3838 let mut build = Builder::new(&mut func, block);
3839 build.store(args[0], args[0], plain(), Flags::default());
3840 build.ret(&[]);
3841
3842 // `void f(void **p)`. It takes a pointer first and returns nothing, which is the shape of
3843 // the one above and none of its meaning, and what tells them apart is the signature. A
3844 // `void` function leaves `rax` alone.
3845 assert!(!lower(&mut names, &func).contains("ret_val"));
3846 }
3847
3848 #[test]
3849 fn a_return_of_a_constant_puts_it_in_a_register_first() {
3850 let (mut names, mut func, block, _) = blank(&[]);
3851 let mut build = Builder::new(&mut func, block);
3852 let zero = build.iconst(Type::int(32), 0);
3853 build.ret(&[zero]);
3854
3855 // No rule returns an immediate, so the plan that offers one is turned down and the next
3856 // one materializes it. That is `int main(void) { return 0; }` in full, once the epilogue
3857 // is appended to it.
3858 assert_eq!(
3859 lower(&mut names, &func),
3860 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_ri_32 0\n x64.ret_val_32 %0($rax)\n}\n"
3861 );
3862 }
3863
3864 #[test]
3865 fn the_rule_that_writes_a_constant_down_is_recorded_as_a_rule_that_fired() {
3866 let (mut names, mut func, block, _) = blank(&[]);
3867 let mut build = Builder::new(&mut func, block);
3868 let zero = build.iconst(Type::int(32), 0);
3869 build.ret(&[zero]);
3870
3871 // The loop over the instructions passes a constant by, because a constant is written where
3872 // a register for it is first wanted rather than where the IR put it. So the only place a
3873 // rule about one is ever selected is the materialization, and a mark made in the loop
3874 // alone would report every rule about a constant as a rule nothing reaches.
3875 let out = super::func(&func, &mut names, &SYSV, &Elsewhere::default())
3876 .expect("every instruction has a rule");
3877 let rules = &crate::select::x86_64::TABLE.rules;
3878 let fired: Vec<&str> = rules
3879 .iter()
3880 .enumerate()
3881 .filter(|(index, _)| out.fired.has(*index))
3882 .map(|(_, rule)| rule.pattern)
3883 .collect();
3884 assert!(fired.contains(&"(iconst.i32 k)"), "{fired:?}");
3885 }
3886
3887 #[test]
3888 fn a_return_of_nothing_is_no_instruction_at_all() {
3889 let (mut names, mut func, block, _) = blank(&[]);
3890 let mut build = Builder::new(&mut func, block);
3891 build.ret(&[]);
3892
3893 // Every part of leaving a function that returns nothing is the epilogue's, and the
3894 // epilogue goes in after allocation. A block with nothing in it is the right answer here
3895 // rather than a function that could not be lowered.
3896 assert_eq!(lower(&mut names, &func), "mfunc @f {\nblock0:\n}\n");
3897 }
3898
3899 #[test]
3900 fn the_allocator_is_what_moves_the_answer_into_the_return_register() {
3901 let (mut names, mut source, block, _) = blank(&[]);
3902 let mut build = Builder::new(&mut source, block);
3903 let zero = build.iconst(Type::int(32), 0);
3904 build.ret(&[zero]);
3905
3906 let mut out = func(&source, &mut names, &SYSV, &Elsewhere::default())
3907 .expect("every instruction has a rule")
3908 .func;
3909 let env = env();
3910 let allocation = rucc_regalloc::run(&mut out, &env, "test");
3911 let frame = Frame::of(&out, &allocation, &Layout::new(&SYSV, REGS));
3912 finish(
3913 &mut out,
3914 &allocation,
3915 &frame,
3916 &Stack::default(),
3917 Convention::new(&SYSV, &FRAME),
3918 &mut names,
3919 );
3920
3921 // `int main(void) { return 0; }` end to end. Nothing here asked for `rax`: the rule said
3922 // the value goes back, the target said where, and the allocator is what made it true. The
3923 // epilogue is what leaves, and this function needs no frame, so it is the return alone.
3924 //
3925 // Two instructions and no copy, which is what a hint buys. The return insists on `rax`,
3926 // so `rax` is the register the allocator tries first for the value the return reads, and
3927 // the constant is written straight into it.
3928 assert_eq!(
3929 mir::print_func(&out, &names, ®S),
3930 "mfunc @f {\nblock0:\n $rax = x64.mov_ri_32 0\n \
3931 x64.ret_val_32 $rax($rax)\n x64.ret\n}\n"
3932 );
3933 }
3934
3935 #[test]
3936 fn a_function_of_two_arguments_is_a_whole_function_now() {
3937 let i32 = Type::int(32);
3938 let (mut names, mut source, block, args) = blank(&[i32, i32]);
3939 let mut build = Builder::new(&mut source, block);
3940 let sum = build.binary(Opcode::Add, args[0], args[1], Flags::default());
3941 build.ret(&[sum]);
3942
3943 let mut out = func(&source, &mut names, &SYSV, &Elsewhere::default())
3944 .expect("every instruction has a rule")
3945 .func;
3946 let env = env();
3947 let allocation = rucc_regalloc::run(&mut out, &env, "test");
3948 let frame = Frame::of(&out, &allocation, &Layout::new(&SYSV, REGS));
3949 finish(
3950 &mut out,
3951 &allocation,
3952 &frame,
3953 &Stack::default(),
3954 Convention::new(&SYSV, &FRAME),
3955 &mut names,
3956 );
3957
3958 // `int f(int a, int b) { return a + b; }` end to end, and this is the test the argument
3959 // side exists for. Before it there was no way to write one: the allocator refuses a
3960 // function whose entry block takes parameters, because there is no edge into an entry
3961 // block for the moves that give a block parameter its value to go on.
3962 //
3963 // One move, and it is the one the machine's addition needs rather than one the allocator
3964 // owes anybody. Each argument stays in the register it arrived in, because the pseudo
3965 // that defines it insists on that register and the allocator now tries it first, and the
3966 // sum stays in the register the addition wrote it to until the return reads it out. The
3967 // copy in front of a two address instruction is what makes its destination one of the
3968 // registers it reads, and the source operand keeps its own name because the destination
3969 // is what the encoder writes.
3970 assert_eq!(
3971 mir::print_func(&out, &names, ®S),
3972 "mfunc @f {\nblock0:\n $rdi($rdi) = x64.arg_val_32\n \
3973 $rsi($rsi) = x64.arg_val_32\n \
3974 $rdi(reuse 1) = x64.add_rr_32 $rdi, $rsi\n $rax = x64.mov_rr_64 $rdi\n \
3975 x64.ret_val_32 $rax($rax)\n x64.ret\n}\n"
3976 );
3977 }
3978
3979 #[test]
3980 fn an_argument_with_no_register_left_for_it_is_read_out_of_the_caller_s_stack() {
3981 let i64 = Type::int(64);
3982 let (mut names, mut source, block, args) = blank(&[i64; 7]);
3983 let mut build = Builder::new(&mut source, block);
3984 build.ret(&[args[6]]);
3985
3986 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
3987 .expect("the seventh is read from memory");
3988
3989 // SysV passes six integers in registers and the seventh in the caller's memory, so six of
3990 // these are pseudos that encode to nothing and the seventh is a load that encodes to real
3991 // bytes. Its displacement is nothing here for the reason a local's is: there is no frame
3992 // yet. What the walk hands on is which instruction is waiting, and for how far up the
3993 // caller's argument area, which is the bottom of it because it is the first one there.
3994 assert_eq!(lowered.stack.arguments.len(), 1);
3995 assert_eq!(lowered.stack.arguments[0].1, 0);
3996 let text = mir::print_func(&lowered.func, &names, ®S);
3997 assert!(text.contains("%6:gpr = x64.mov_rm_64 [$rsp]"), "{text}");
3998 assert_eq!(text.matches("x64.arg_val_64").count(), 6, "{text}");
3999 }
4000
4001 #[test]
4002 fn the_frame_is_what_says_how_far_up_the_caller_s_stack_an_argument_is() {
4003 let i64 = Type::int(64);
4004 let (mut names, mut source, block, args) = blank(&[i64; 8]);
4005 let mut build = Builder::new(&mut source, block);
4006 let sum = build.binary(Opcode::Add, args[6], args[7], Flags::default());
4007 build.ret(&[sum]);
4008
4009 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4010 .expect("both are read from memory");
4011 let stack = lowered.stack;
4012 let mut out = lowered.func;
4013 let env = env();
4014 let allocation = rucc_regalloc::run(&mut out, &env, "test");
4015 let layout = stack.layout(Layout::new(&SYSV, REGS));
4016 let frame = Frame::of(&out, &allocation, &layout);
4017 finish(&mut out, &allocation, &frame, &stack, Convention::new(&SYSV, &FRAME), &mut names);
4018
4019 // A leaf that takes no frame, so the stack pointer never moves and the only thing between
4020 // it and the caller's arguments is the return address the call pushed. The seventh
4021 // parameter is at the bottom of the caller's argument area and the eighth is one word
4022 // further up, which is the eight bytes between the two offsets.
4023 let text = mir::print_func(&out, &names, ®S);
4024 assert_eq!(frame.size(), 0);
4025 assert_eq!(frame.incoming(), Incoming::from_stack(8));
4026 assert!(text.contains("x64.mov_rm_64 [$rsp + 8]"), "{text}");
4027 assert!(text.contains("x64.mov_rm_64 [$rsp + 16]"), "{text}");
4028 }
4029
4030 #[test]
4031 fn a_realigned_frame_reaches_the_caller_s_arguments_through_the_frame_pointer() {
4032 let i64 = Type::int(64);
4033 let (mut names, mut source, block, args) = blank(&[i64; 7]);
4034 let wide = slot(&mut source, block, 64, 32);
4035 let mut build = Builder::new(&mut source, block);
4036 build.store(args[6], wide, plain(), Flags::default());
4037 build.ret(&[args[6]]);
4038
4039 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4040 .expect("every instruction has a rule");
4041 let stack = lowered.stack;
4042 let mut out = lowered.func;
4043 let env = env();
4044 let allocation = rucc_regalloc::run(&mut out, &env, "test");
4045 let layout = stack.layout(Layout::new(&SYSV, REGS));
4046 let frame = Frame::of(&out, &allocation, &layout);
4047 finish(&mut out, &allocation, &frame, &stack, Convention::new(&SYSV, &FRAME), &mut names);
4048
4049 // A local wanting thirty two byte alignment makes the prologue force the stack pointer,
4050 // which throws away how far the caller's stack was. So the load the lowering wrote off the
4051 // stack pointer is rewritten to read through the frame pointer, at the one distance that
4052 // survives: the word the prologue pushed the frame pointer into, and the return address
4053 // above it.
4054 let text = mir::print_func(&out, &names, ®S);
4055 assert_eq!(frame.realign(), Some(32));
4056 assert_eq!(frame.incoming(), Incoming::from_frame(16));
4057 assert!(text.contains("x64.mov_rm_64 [$rbp + 16]"), "{text}");
4058 assert!(!text.contains("x64.mov_rm_64 [$rsp"), "{text}");
4059 }
4060
4061 #[test]
4062 fn a_jump_is_the_edge_and_nothing_else() {
4063 let i32 = Type::int(32);
4064 let (mut names, mut source, entry, args) = blank(&[i32]);
4065 let next = source.create_block();
4066 let got = source.append_param(next, i32);
4067 Builder::new(&mut source, entry).jump(next, &[args[0]]);
4068 Builder::new(&mut source, next).ret(&[got]);
4069
4070 // Two blocks and two instructions, and the jump is neither of them. What it was is the
4071 // arm on the first block, and what the arm carries is the argument it was called with.
4072 assert_eq!(
4073 lower(&mut names, &source),
4074 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32 block1(%0)\n\n\
4075 block1(%1:gpr):\n x64.ret_val_32 %1($rax)\n}\n"
4076 );
4077 }
4078
4079 /// A block that reads what a block below it writes is filled after it, not before it.
4080 ///
4081 /// The blocks are written entry, `early`, `late`, `exit`, and the entry jumps straight past
4082 /// `early` to `late`, so `late` dominates `early` while sitting below it in the function.
4083 /// Filling them in the order they are written reaches the read in `early` first, and reading
4084 /// a value with no register yet mints one. The cast in `late` is no instruction at all, so
4085 /// what it does is give its answer the register its operand is already in, and that is not
4086 /// the register the read minted. Nothing writes the register the read minted. The printer
4087 /// says `%?` for a register nothing defines, which is what this looks for, and what came out
4088 /// of the real bug was SQLite loading a stack slot no store ever reached.
4089 #[test]
4090 fn a_block_that_reads_what_a_block_below_it_writes_is_filled_after_it() {
4091 let i64 = Type::int(64);
4092 let (mut names, mut source, entry, args) = blank(&[i64, i64]);
4093 let early = source.create_block();
4094 let late = source.create_block();
4095 let exit = source.create_block();
4096
4097 Builder::new(&mut source, entry).jump(late, &[]);
4098 let ptr = cast(&mut source, late, Opcode::IntToPtr, args[0], Type::PTR);
4099 Builder::new(&mut source, early).ret(&[ptr]);
4100 let mut build = Builder::new(&mut source, late);
4101 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
4102 build.br_if(cond, early, &[], exit, &[]);
4103 Builder::new(&mut source, exit).ret(&[args[1]]);
4104
4105 let text = lower(&mut names, &source);
4106 assert!(!text.contains("%?"), "every register has something that writes it: {text}");
4107 }
4108
4109 /// A constant is written where it is wanted rather than where the IR defined it, and two
4110 /// blocks wanting the same one is two places. Writing it once and reading it in both is a
4111 /// register read where nothing wrote it, unless the block it was written in happens to
4112 /// dominate the other, which nothing here checks and which the second arm of a branch never
4113 /// does. Each block gets its own copy of the number instead.
4114 #[test]
4115 fn a_constant_two_blocks_want_is_written_in_both_of_them() {
4116 let i32 = Type::int(32);
4117 let (mut names, mut source, entry, args) = blank(&[i32, i32]);
4118 let then = source.create_block();
4119 let other = source.create_block();
4120 let join = source.create_block();
4121 let got = source.append_param(join, i32);
4122
4123 let mut build = Builder::new(&mut source, entry);
4124 let seven = build.iconst(i32, 7);
4125 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
4126 build.br_if(cond, then, &[], other, &[]);
4127 // Both arms want the seven in a register, because a block argument is never an immediate,
4128 // and neither arm dominates the other.
4129 Builder::new(&mut source, then).jump(join, &[seven]);
4130 Builder::new(&mut source, other).jump(join, &[seven]);
4131 Builder::new(&mut source, join).ret(&[got]);
4132
4133 let text = lower(&mut names, &source);
4134 assert_eq!(text.matches("x64.mov_ri_32 7").count(), 2, "one seven per block: {text}");
4135 }
4136
4137 /// An argument on an edge out of a block that leaves two ways is read after every instruction
4138 /// of the block is written, and reading one can write an instruction, which would land after
4139 /// the branch that has already jumped past it. The branch goes back on the end.
4140 #[test]
4141 fn a_constant_an_edge_wants_is_written_before_the_branch_and_not_after_it() {
4142 let i32 = Type::int(32);
4143 let (mut names, mut source, entry, args) = blank(&[i32, i32]);
4144 let then = source.create_block();
4145 let join = source.create_block();
4146 let got = source.append_param(join, i32);
4147
4148 let mut build = Builder::new(&mut source, entry);
4149 let nine = build.iconst(i32, 9);
4150 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
4151 build.br_if(cond, then, &[], join, &[nine]);
4152 Builder::new(&mut source, then).jump(join, &[args[0]]);
4153 Builder::new(&mut source, join).ret(&[got]);
4154
4155 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
4156 .expect("every instruction has a rule")
4157 .func;
4158 let entry = out.entry().expect("an entry block");
4159 let last = out.terminator(entry).expect("a block that leaves two ways has a branch");
4160 let branch = names.intern("x64.br_cond_8");
4161 assert_eq!(
4162 out[last].opcode,
4163 mir::Opcode::new(branch),
4164 "the branch is last: {}",
4165 mir::print_func(&out, &names, ®S)
4166 );
4167 }
4168
4169 #[test]
4170 fn a_conditional_branch_is_lowered_to_the_condition_and_nothing_about_where_it_goes() {
4171 let i32 = Type::int(32);
4172 let (mut names, mut source, entry, args) = blank(&[i32, i32]);
4173 let then = source.create_block();
4174 let other = source.create_block();
4175 let mut build = Builder::new(&mut source, entry);
4176 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
4177 build.br_if(cond, then, &[], other, &[]);
4178 Builder::new(&mut source, then).ret(&[args[0]]);
4179 Builder::new(&mut source, other).ret(&[args[1]]);
4180
4181 // The comparison writes a byte and the branch reads it, and neither says a block. Both
4182 // arms are on the entry block, in the order the branch took them, so the arm that runs
4183 // when the condition holds is the first.
4184 assert_eq!(
4185 lower(&mut names, &source),
4186 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
4187 %1:gpr($rsi) = x64.arg_val_32\n %2:gpr = x64.cmp_set_l_32 %0, %1\n \
4188 x64.br_cond_8 %2, block1, block2\n\n\
4189 block1:\n x64.ret_val_32 %0($rax)\n\n\
4190 block2:\n x64.ret_val_32 %1($rax)\n}\n"
4191 );
4192 }
4193
4194 /// A choice between two values, which is one instruction and no blocks at all.
4195 ///
4196 /// The arms come out the other way round from the IR, because a conditional move overwrites its
4197 /// destination and the destination is the arm taken when the condition does not hold. The
4198 /// condition arrives last for the same reason: it is read by the test in front of the move
4199 /// rather than by the move.
4200 #[test]
4201 fn a_select_is_lowered_to_a_test_and_a_conditional_move() {
4202 let i32 = Type::int(32);
4203 let (mut names, mut source, entry, args) = blank(&[i32, i32]);
4204 let mut build = Builder::new(&mut source, entry);
4205 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
4206 let picked = build.select(cond, args[0], args[1]);
4207 build.ret(&[picked]);
4208
4209 assert_eq!(
4210 lower(&mut names, &source),
4211 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
4212 %1:gpr($rsi) = x64.arg_val_32\n %2:gpr = x64.cmp_set_l_32 %0, %1\n \
4213 %3:gpr(reuse 1) = x64.test_cmov_ne_32 %1, %0, %2\n \
4214 x64.ret_val_32 %3($rax)\n}\n"
4215 );
4216 }
4217
4218 #[test]
4219 fn a_branch_over_a_block_is_a_whole_function_now() {
4220 let i32 = Type::int(32);
4221 let (mut names, mut source, entry, args) = blank(&[i32, i32]);
4222 let then = source.create_block();
4223 let other = source.create_block();
4224 let join = source.create_block();
4225 let got = source.append_param(join, i32);
4226 let mut build = Builder::new(&mut source, entry);
4227 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
4228 build.br_if(cond, then, &[], other, &[]);
4229 let mut build = Builder::new(&mut source, then);
4230 let sum = build.binary(Opcode::Add, args[0], args[1], Flags::default());
4231 build.jump(join, &[sum]);
4232 Builder::new(&mut source, other).jump(join, &[args[1]]);
4233 Builder::new(&mut source, join).ret(&[got]);
4234
4235 // `int f(int a, int b) { if (a < b) return a + b; else return b; }` end to end, written
4236 // the way a front end writes it: both arms of the branch are blocks of their own and the
4237 // return is the block they meet at. No edge here is critical, because the two arms out of
4238 // the entry carry nothing and the two arms into the join each leave a block that goes
4239 // nowhere else, so each has its own end to put its move at.
4240 let mut out = func(&source, &mut names, &SYSV, &Elsewhere::default())
4241 .expect("every instruction has a rule")
4242 .func;
4243 assert_eq!(crate::split::critical(&mut out), 0, "no edge here is critical");
4244 let env = env();
4245 let allocation = rucc_regalloc::run(&mut out, &env, "test");
4246 let frame = Frame::of(&out, &allocation, &Layout::new(&SYSV, REGS));
4247 finish(
4248 &mut out,
4249 &allocation,
4250 &frame,
4251 &Stack::default(),
4252 Convention::new(&SYSV, &FRAME),
4253 &mut names,
4254 );
4255
4256 // One epilogue, on the join, which is the one block the function leaves from, and the
4257 // moves that give the join its parameter are at the end of each arm. Every register is
4258 // physical and the branch is still a branch on a register, because turning it into a
4259 // `test` and a `jcc` is the block layout's and there is no block layout yet.
4260 let text = mir::print_func(&out, &names, ®S);
4261 assert_eq!(text.matches("x64.ret\n").count(), 1, "{text}");
4262 assert!(text.contains("x64.br_cond_8"), "{text}");
4263 assert!(text.contains("x64.add_rr_32"), "{text}");
4264 assert!(!text.contains('%'), "{text}");
4265 }
4266
4267 #[test]
4268 fn a_critical_edge_is_split_before_the_allocator_ever_sees_it() {
4269 let i32 = Type::int(32);
4270 let (mut names, mut source, entry, args) = blank(&[i32, i32]);
4271 let then = source.create_block();
4272 let join = source.create_block();
4273 let got = source.append_param(join, i32);
4274 let mut build = Builder::new(&mut source, entry);
4275 let cond = build.icmp(rucc_ir::IntPred::Slt, args[0], args[1]);
4276 build.br_if(cond, then, &[], join, &[args[1]]);
4277 Builder::new(&mut source, then).jump(join, &[args[0]]);
4278 let mut build = Builder::new(&mut source, join);
4279 let twice = build.binary(Opcode::Add, got, got, Flags::default());
4280 build.ret(&[twice]);
4281
4282 // The else arm is critical: the entry block leaves two ways and the join is arrived at
4283 // two ways, and the arm carries a value. Without splitting it the allocator asserts,
4284 // because the move that gives the join its parameter would have to run at the end of a
4285 // block that also goes to the other arm.
4286 let mut out = func(&source, &mut names, &SYSV, &Elsewhere::default())
4287 .expect("every instruction has a rule")
4288 .func;
4289 assert_eq!(crate::split::critical(&mut out), 1);
4290 let env = env();
4291 let allocation = rucc_regalloc::run(&mut out, &env, "test");
4292 let frame = Frame::of(&out, &allocation, &Layout::new(&SYSV, REGS));
4293 finish(
4294 &mut out,
4295 &allocation,
4296 &frame,
4297 &Stack::default(),
4298 Convention::new(&SYSV, &FRAME),
4299 &mut names,
4300 );
4301
4302 // The block the split added is where the move went, and it is the whole of that block.
4303 let text = mir::print_func(&out, &names, ®S);
4304 assert_eq!(out.block_count(), 4, "{text}");
4305 assert_eq!(text.matches("x64.ret\n").count(), 1, "{text}");
4306 }
4307
4308 #[test]
4309 fn a_call_passes_what_the_convention_says_and_takes_back_what_it_says() {
4310 let i32 = Type::int(32);
4311 let (mut names, mut source, block, args) = blank(&[i32, i32]);
4312 let sig =
4313 source.add_signature(Signature::new().with_params(&[i32, i32]).with_returns(&[i32]));
4314 let callee = names.intern("g");
4315 let call = Builder::new(&mut source, block).call(callee, sig, &[args[0], args[1]]);
4316 let got = source[call].first_result.expect("an integer comes back");
4317 Builder::new(&mut source, block).ret(&[got]);
4318
4319 // `int f(int a, int b) { return g(a, b); }`. The arguments arrived where the call wants
4320 // them, so what the call reads is what arrived, and the whole of the convention is in the
4321 // constraints rather than in a move.
4322 let text = lower(&mut names, &source);
4323 assert!(text.contains("= x64.call %0($rdi), %1($rsi), @g"), "{text}");
4324 assert!(text.contains("x64.ret_val_32 %2($rax)"), "{text}");
4325 // What the call writes is the value that comes back and then every register the callee is
4326 // free to destroy, in both classes, which is the whole of what stops the allocator from
4327 // leaving something in one of them.
4328 assert!(text.contains("%2:gpr($rax), $rcx, $rdx, $r8, $r9, $r10, $r11, $xmm0,"), "{text}");
4329 assert!(text.contains("$xmm15 = x64.call"), "{text}");
4330 }
4331
4332 #[test]
4333 fn what_the_frame_owes_a_call_comes_back_with_the_function() {
4334 let i32 = Type::int(32);
4335 let sig = |source: &mut Func| source.add_signature(Signature::new().with_params(&[i32]));
4336
4337 let (mut names, mut source, block, args) = blank(&[i32]);
4338 let sig = sig(&mut source);
4339 let callee = names.intern("g");
4340 Builder::new(&mut source, block).call(callee, sig, &[args[0]]);
4341 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
4342 .expect("every instruction has a rule");
4343
4344 // Nothing on the stack, so nothing owed, but not a leaf either: a function that calls
4345 // owes the callee an aligned stack pointer and may not use the red zone.
4346 assert_eq!(out.stack.calls, Some(0));
4347 let layout = out.stack.layout(Layout::new(&SYSV, REGS));
4348 assert!(!layout.leaf);
4349 assert_eq!(layout.outgoing, 0);
4350
4351 // The same call under the other convention owes thirty two bytes for the callee to spill
4352 // its register arguments into, which is a fact about the convention and not about the call.
4353 let out = func(&source, &mut names, &x86_64::WIN64, &Elsewhere::default())
4354 .expect("every instruction has a rule");
4355 assert_eq!(out.stack.calls, Some(32));
4356
4357 // And a function that calls nothing is a leaf, which is what says it may use the red zone.
4358 let (mut names, mut source, block, args) = blank(&[i32]);
4359 Builder::new(&mut source, block).ret(&[args[0]]);
4360 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
4361 .expect("every instruction has a rule");
4362 assert_eq!(out.stack.calls, None);
4363 assert!(out.stack.layout(Layout::new(&SYSV, REGS)).leaf);
4364 }
4365
4366 #[test]
4367 fn a_value_that_outlives_a_call_is_not_left_where_the_call_destroys_it() {
4368 let i32 = Type::int(32);
4369 let (mut names, mut source, block, args) = blank(&[i32]);
4370 let sig = source.add_signature(Signature::new().with_params(&[i32]).with_returns(&[i32]));
4371 let callee = names.intern("g");
4372 let call = Builder::new(&mut source, block).call(callee, sig, &[args[0]]);
4373 let got = source[call].first_result.expect("an integer comes back");
4374 let mut build = Builder::new(&mut source, block);
4375 let sum = build.binary(Opcode::Add, got, args[0], Flags::default());
4376 build.ret(&[sum]);
4377
4378 // `int f(int a) { return g(a) + a; }`, which is the smallest program that asks the
4379 // question: `a` is read after the call and `rdi` is a register the call destroys.
4380 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4381 .expect("every instruction has a rule");
4382 let layout = lowered.stack.layout(Layout::new(&SYSV, REGS));
4383 let mut out = lowered.func;
4384 let env = env();
4385 let allocation = rucc_regalloc::run(&mut out, &env, "test");
4386 let frame = Frame::of(&out, &allocation, &layout);
4387 finish(
4388 &mut out,
4389 &allocation,
4390 &frame,
4391 &Stack::default(),
4392 Convention::new(&SYSV, &FRAME),
4393 &mut names,
4394 );
4395
4396 // It went to a register the callee has to put back, and the prologue and epilogue are what
4397 // put it back, which is the whole bargain the two halves of a convention make.
4398 let text = mir::print_func(&out, &names, ®S);
4399 assert!(text.contains("$rbx"), "{text}");
4400 assert!(!text.contains('%'), "{text}");
4401 assert_eq!(text.matches("x64.call").count(), 1, "{text}");
4402 }
4403
4404 #[test]
4405 fn a_call_with_more_arguments_than_registers_writes_the_rest_into_the_outgoing_area() {
4406 let i64 = Type::int(64);
4407 let (mut names, mut source, block, args) = blank(&[i64]);
4408 let seven = vec![i64; 7];
4409 let sig = source.add_signature(Signature::new().with_params(&seven));
4410 let callee = names.intern("g");
4411 let passed = vec![args[0]; 7];
4412 Builder::new(&mut source, block).call(callee, sig, &passed);
4413
4414 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4415 .expect("the seventh goes to memory");
4416 // The bytes the call needs are on the layout the frame is worked out from, so that the
4417 // frame reserves as many as the widest call in the function asked for.
4418 assert_eq!(lowered.stack.calls, Some(8));
4419 let text = mir::print_func(&lowered.func, &names, ®S);
4420 assert!(text.contains("x64.mov_mr_64 %0, [$rsp]\n"), "{text}");
4421 }
4422
4423 #[test]
4424 fn a_call_this_cannot_make_is_reported_rather_than_made() {
4425 let (mut names, mut source, block, _) = blank(&[]);
4426 let returns = [Type::float(rucc_ir::Float::F80), Type::int(64)];
4427 let sig = source.add_signature(Signature::new().with_returns(&returns));
4428 let callee = names.intern("g");
4429 Builder::new(&mut source, block).call(callee, sig, &[]);
4430 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4431 .expect_err("a long double is on the x87");
4432 assert_eq!(failed.to_string(), "what this call gives back is on the x87 stack");
4433 }
4434
4435 /// A `long double` on its own is a different answer, because on its own it comes back on the
4436 /// x87 stack rather than in a register, which is somewhere the call cannot be said to write.
4437 ///
4438 /// So the call gives back nothing at all and the value is taken off the stack by the `fstp`
4439 /// straight after it. That instruction has to be straight after it: the stack is one place and
4440 /// anything else that touched it before this ran would be looking at the value still on it.
4441 #[test]
4442 fn a_call_that_gives_back_a_long_double_takes_it_off_the_stack_at_once() {
4443 let (mut names, mut source, block, _) = blank(&[]);
4444 let long_double = Type::float(rucc_ir::Float::F80);
4445 let sig = source.add_signature(Signature::new().with_returns(&[long_double]));
4446 let callee = names.intern("g");
4447 Builder::new(&mut source, block).call(callee, sig, &[]);
4448
4449 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4450 .expect("the value comes back in st0");
4451 let text = mir::print_func(&lowered.func, &names, ®S);
4452 let after: Vec<&str> =
4453 text.lines().skip_while(|line| !line.contains("x64.call")).skip(1).collect();
4454 assert_eq!(after[0].trim(), "%0:gpr = x64.lea_64 [$rsp]", "{text}");
4455 assert_eq!(after[1].trim(), "x64.fstp_t [%0]", "{text}");
4456 // And the slot it went into is the sixteen bytes the type takes, like every other one.
4457 assert_eq!(lowered.stack.locals.len(), 1, "{text}");
4458 assert_eq!(lowered.stack.locals[0].size, X87_BYTES);
4459 }
4460
4461 #[test]
4462 fn a_call_through_an_address_goes_through_the_register_the_address_is_in() {
4463 let i32 = Type::int(32);
4464 let (mut names, mut source, block, args) = blank(&[Type::PTR, i32]);
4465 let sig = source.add_signature(Signature::new().with_params(&[i32]).with_returns(&[i32]));
4466 let varargs = source.push_abis(&[]);
4467 let info = source.add_call(CallInfo { callee: None, signature: sig, varargs });
4468 let mut build = Builder::new(&mut source, block);
4469 let inst = InstData {
4470 args: build.func().push_values(&[args[0], args[1]]),
4471 extra: Extra::Call(info),
4472 ..InstData::new(Opcode::CallIndirect)
4473 };
4474 let called = build.inst(inst, &[i32]);
4475 let got = source[called].first_result.expect("an integer comes back");
4476 Builder::new(&mut source, block).ret(&[got]);
4477
4478 // `int f(int (*g)(int), int a) { return g(a); }`. The first operand is the address and
4479 // the arguments are the ones behind it, and everything else about the call is what a call
4480 // to a name would have been.
4481 let text = lower(&mut names, &source);
4482 assert!(text.contains("= x64.call_reg %0, %1($rdi)"), "{text}");
4483 assert!(text.contains("x64.ret_val_32 %2($rax)"), "{text}");
4484 assert!(!text.contains("@g"), "a call through an address names nobody: {text}");
4485 }
4486
4487 #[test]
4488 fn an_instruction_no_rule_covers_is_reported() {
4489 let (mut names, mut source, block, args) = blank(&[Type::PTR]);
4490 let mut build = Builder::new(&mut source, block);
4491 let operands = build.func().push_values(&[args[0]]);
4492 build.inst(InstData { args: operands, ..InstData::new(Opcode::LongjmpMarker) }, &[]);
4493
4494 // The mark that a jump goes back through here, which nothing writes an instruction for
4495 // yet: what it needs is for the allocator to be told a block can be arrived at twice, and
4496 // that is `tamnd/rucc#223`. Nothing about it is a width or a register, so there is nothing
4497 // for the message to add beyond the name.
4498 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4499 .expect_err("no rule writes a longjmp marker");
4500 assert_eq!(failed.to_string(), "no rule lowers a `longjmp_marker`");
4501
4502 // It produces nothing, so there is no type in the message and nothing invents one, and the
4503 // instruction comes back so a caller can ask the function where it was.
4504 let inst = failed.inst().expect("the instruction it is about");
4505 assert_eq!(source[inst].opcode, Opcode::LongjmpMarker);
4506 }
4507
4508 /// A barrier is written by name here, and what it is depends on the ordering and on nothing
4509 /// else. `crate::expand` is where the reasoning about this machine's memory model lives.
4510 #[test]
4511 fn a_barrier_is_one_instruction_at_the_strongest_ordering_and_none_below_it() {
4512 for order in MemOrder::all().filter(|&order| order != MemOrder::NotAtomic) {
4513 let (mut names, mut source, block, _) = blank(&[]);
4514 let mut build = Builder::new(&mut source, block);
4515 build
4516 .inst(InstData { extra: Extra::Order(order), ..InstData::new(Opcode::Fence) }, &[]);
4517
4518 let text = lower(&mut names, &source);
4519 assert_eq!(text.contains("x64.mfence"), order == MemOrder::SeqCst, "{order:?}: {text}");
4520 }
4521 }
4522
4523 /// A compare and exchange is written by name too, and at the width of the value rather than at
4524 /// the width of the address, which is the mistake worth pinning: everything here is a pointer
4525 /// and only the value says how many bytes the instruction touches.
4526 #[test]
4527 fn a_compare_and_exchange_is_one_instruction_at_the_width_of_the_value() {
4528 for bits in [8, 16, 32, 64] {
4529 let ty = Type::int(bits);
4530 let (mut names, mut source, block, args) = blank(&[Type::PTR, ty, ty]);
4531 let mut build = Builder::new(&mut source, block);
4532 let mem = build.func().add_mem(MemInfo {
4533 size: u64::from(bits / 8),
4534 align: bits / 8,
4535 order: MemOrder::SeqCst,
4536 ..plain()
4537 });
4538 let operands = build.func().push_values(&[args[0], args[1], args[2]]);
4539 build.inst(
4540 InstData {
4541 args: operands,
4542 extra: Extra::Mem(mem),
4543 ..InstData::new(Opcode::Cmpxchg)
4544 },
4545 &[ty, Type::I1],
4546 );
4547
4548 // Two values out of one instruction, the first of them in the register the machine
4549 // reads the expected value out of, the second free for the allocator to place. The
4550 // address is the memory operand and neither of the two values is.
4551 let text = lower(&mut names, &source);
4552 let written = format!("%3:gpr($rax), %4:gpr = x64.cmpxchg_{bits} %1($rax), %2, [%0]");
4553 assert!(text.contains(&written), "{bits}: {text}");
4554 }
4555 }
4556
4557 #[test]
4558 fn more_values_back_than_the_convention_has_registers_for_is_reported() {
4559 let i64 = Type::int(64);
4560 let (mut names, mut source, block, args) = blank(&[i64, i64, i64]);
4561 let mut build = Builder::new(&mut source, block);
4562 build.ret(&[args[0], args[1], args[2]]);
4563
4564 // Two integers come back in `rax` and `rdx` and a third has nowhere to go, which is not a
4565 // gap in the rules but the convention saying no. The front end classifies before it gets
4566 // here, so this is the shape that would mean the classification went wrong.
4567 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4568 .expect_err("only two come back");
4569 assert_eq!(
4570 failed.to_string(),
4571 "what this function gives back takes more registers than this convention has for it"
4572 );
4573
4574 let inst = failed.inst().expect("the instruction it is about");
4575 assert_eq!(source[inst].opcode, Opcode::Return);
4576 }
4577
4578 /// A refusal about a signature has no instruction, which is what makes it the one arm apart.
4579 ///
4580 /// Everything else is about something written somewhere in the body and hands it back so a
4581 /// caller can ask the function where it came from. A parameter arrives before the first
4582 /// instruction runs, so there is nothing in the body to point at and the message is about
4583 /// the function.
4584 #[test]
4585 fn a_refusal_about_a_parameter_has_no_instruction_to_point_at() {
4586 let missing = Unsupported::Argument { index: 0, missing: Missing::OnX87 };
4587 assert_eq!(missing.inst(), None);
4588 }
4589
4590 /// An `alloca` of a fixed size, which is what every local whose address is taken becomes.
4591 fn slot(source: &mut Func, block: Block, size: u64, align: u32) -> Value {
4592 let info = MemInfo { size, align, ..plain() };
4593 let mut build = Builder::new(source, block);
4594 let mem = build.func().add_mem(info);
4595 build.value(InstData { extra: Extra::Mem(mem), ..InstData::new(Opcode::Alloca) }, Type::PTR)
4596 }
4597
4598 #[test]
4599 fn a_local_is_memory_in_the_frame_and_one_instruction_that_says_where() {
4600 let (mut names, mut source, block, _) = blank(&[]);
4601 let slot = slot(&mut source, block, 4, 4);
4602 let mut build = Builder::new(&mut source, block);
4603 let nine = build.iconst(Type::int(32), 9);
4604 build.store(nine, slot, plain(), Flags::default());
4605 let loaded = build.load(Type::int(32), slot, plain(), Flags::default());
4606 build.ret(&[loaded]);
4607
4608 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4609 .expect("every instruction has a rule");
4610
4611 // Four bytes on the list the frame is laid out from, and the one instruction that reads
4612 // where they went. Its displacement is nothing here because there is no frame yet, and
4613 // which instruction is waiting for which local is what `finish` is handed.
4614 assert_eq!(lowered.stack.locals, vec![Local { size: 4, align: 4 }]);
4615 assert_eq!(lowered.stack.addresses.len(), 1);
4616 assert_eq!(lowered.stack.addresses[0].1, 0);
4617 assert_eq!(
4618 mir::print_func(&lowered.func, &names, ®S),
4619 "mfunc @f {\nblock0:\n %0:gpr = x64.lea_64 [$rsp]\n \
4620 %1:gpr = x64.mov_ri_32 9\n x64.mov_mr_32 %1, [%0]\n \
4621 %2:gpr = x64.mov_rm_32 [%0]\n x64.ret_val_32 %2($rax)\n}\n"
4622 );
4623 }
4624
4625 #[test]
4626 fn the_frame_is_what_fills_the_address_of_a_local_in() {
4627 let (mut names, mut source, block, _) = blank(&[]);
4628 let slot = slot(&mut source, block, 4, 4);
4629 let mut build = Builder::new(&mut source, block);
4630 let nine = build.iconst(Type::int(32), 9);
4631 build.store(nine, slot, plain(), Flags::default());
4632 let loaded = build.load(Type::int(32), slot, plain(), Flags::default());
4633 build.ret(&[loaded]);
4634
4635 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4636 .expect("every instruction has a rule");
4637 let stack = lowered.stack;
4638 let mut out = lowered.func;
4639 let env = env();
4640 let allocation = rucc_regalloc::run(&mut out, &env, "test");
4641 let layout = stack.layout(Layout::new(&SYSV, REGS));
4642 let frame = Frame::of(&out, &allocation, &layout);
4643 finish(&mut out, &allocation, &frame, &stack, Convention::new(&SYSV, &FRAME), &mut names);
4644
4645 // `int f(void) { int x; x = 9; return x; }` with the address of `x` taken, end to end.
4646 // A leaf small enough to live in the red zone takes no frame at all, so the stack pointer
4647 // never moves and the four bytes are below it, which is what the negative offset is. The
4648 // instruction the lowering left with nothing in its displacement now has the answer in it.
4649 let text = mir::print_func(&out, &names, ®S);
4650 assert!(text.contains("$rax = x64.lea_64 [$rsp - 8]"), "{text}");
4651 assert!(!text.contains("x64.sub_ri_64"), "{text}");
4652 assert_eq!(frame.size(), 0);
4653 assert_eq!(frame.local(0), Some(-8));
4654 }
4655
4656 /// An `alloca` whose size is an operand, which is a variable length array.
4657 fn growing(source: &mut Func, block: Block, size: Value, align: u32) -> Value {
4658 let info = MemInfo { size: 0, align, ..plain() };
4659 let mut build = Builder::new(source, block);
4660 let mem = build.func().add_mem(info);
4661 let args = build.func().push_values(&[size]);
4662 build.value(
4663 InstData { args, extra: Extra::Mem(mem), ..InstData::new(Opcode::Alloca) },
4664 Type::PTR,
4665 )
4666 }
4667
4668 #[test]
4669 fn a_stack_slot_whose_size_is_not_known_until_it_runs_takes_the_bytes_off_the_stack_pointer() {
4670 let (mut names, mut source, block, args) = blank(&[Type::int(64)]);
4671 let slot = growing(&mut source, block, args[0], 16);
4672 Builder::new(&mut source, block).ret(&[slot]);
4673
4674 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4675 .expect("every instruction has a rule");
4676
4677 // The bytes come off the stack pointer where the declaration stands and the address is
4678 // where the stack pointer then is, which is one subtraction and one `lea` rather than a
4679 // slot the frame laid out. Nothing is on the list of locals, because there is nothing
4680 // about this the frame could place.
4681 let text = mir::print_func(&lowered.func, &names, ®S);
4682 assert!(text.contains("$rsp = x64.sub_rr_64 $rsp, %0"), "{text}");
4683 assert!(text.contains("x64.lea_64 [$rsp]"), "{text}");
4684 assert!(lowered.stack.locals.is_empty(), "{text}");
4685 assert_eq!(lowered.stack.dynamic.len(), 1);
4686 assert!(lowered.stack.grown_at.is_some());
4687 }
4688
4689 #[test]
4690 fn a_growing_slot_wanting_more_alignment_than_the_stack_pointer_has_is_reported() {
4691 let (mut names, mut source, block, args) = blank(&[Type::int(64)]);
4692 let slot = growing(&mut source, block, args[0], 32);
4693 Builder::new(&mut source, block).ret(&[slot]);
4694
4695 // Thirty two is more than a call leaves the stack pointer on, so giving it what it asked
4696 // for means masking the stack pointer after moving it, and after that no constant reaches
4697 // the rest of the frame from the frame pointer either. A second pointer held for the
4698 // purpose is what fixes it and there is not one yet.
4699 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4700 .expect_err("nothing realigns a frame that grows");
4701 assert_eq!(
4702 failed.to_string(),
4703 "this local wants more alignment than the stack pointer is left on, which needs a \
4704 base register nothing here keeps"
4705 );
4706 }
4707
4708 #[test]
4709 fn a_frame_that_grows_reaches_its_own_locals_through_the_frame_pointer() {
4710 let (mut names, mut source, block, args) = blank(&[Type::int(64)]);
4711 let fixed = slot(&mut source, block, 4, 4);
4712 let mut build = Builder::new(&mut source, block);
4713 let nine = build.iconst(Type::int(32), 9);
4714 build.store(nine, fixed, plain(), Flags::default());
4715 let grown = growing(&mut source, block, args[0], 16);
4716 Builder::new(&mut source, block).ret(&[grown]);
4717
4718 let lowered = func(&source, &mut names, &SYSV, &Elsewhere::default())
4719 .expect("every instruction has a rule");
4720 let stack = lowered.stack;
4721 let mut out = lowered.func;
4722 let env = env();
4723 let allocation = rucc_regalloc::run(&mut out, &env, "test");
4724 let layout = stack.layout(Layout::new(&SYSV, REGS));
4725 let frame = Frame::of(&out, &allocation, &layout);
4726 finish(&mut out, &allocation, &frame, &stack, Convention::new(&SYSV, &FRAME), &mut names);
4727
4728 // The stack pointer moves in the middle of the function, so the four bytes of the fixed
4729 // local are not a constant away from it any more and the frame pointer is what reaches
4730 // them. The frame keeps one whatever the flags asked for, takes its bytes rather than
4731 // living in the red zone, and the address of the growing slot is off the stack pointer as
4732 // it stands after the subtraction rather than off anything the prologue left.
4733 let text = mir::print_func(&out, &names, ®S);
4734 assert!(frame.grows());
4735 assert!(frame.frame_pointer());
4736 assert!(frame.size() > 0, "{text}");
4737 assert!(text.contains("x64.lea_64 [$rbp"), "{text}");
4738 assert!(text.contains("$rsp = x64.sub_rr_64 $rsp"), "{text}");
4739 assert!(text.contains("x64.lea_64 [$rsp]"), "{text}");
4740 }
4741
4742 #[test]
4743 fn an_address_is_read_written_and_added_to_like_the_integer_it_is() {
4744 let (mut names, mut source, block, args) = blank(&[Type::PTR, Type::int(64)]);
4745 let mut build = Builder::new(&mut source, block);
4746 let stepped = build.func().push_values(&[args[0], args[1]]);
4747 let next =
4748 build.value(InstData { args: stepped, ..InstData::new(Opcode::PtrAdd) }, Type::PTR);
4749 let loaded = build.load(Type::int(32), next, plain(), Flags::default());
4750 build.ret(&[loaded]);
4751
4752 // `int f(int *p, long i) { return *(int *)((char *)p + i); }`. Nothing about this is new
4753 // in the rule set, which is the point: the two addresses arrive in registers because an
4754 // address is an integer as wide as one, and the arithmetic on them is the add it always
4755 // was, so every rule written about an add reaches it.
4756 //
4757 // The add stays its own instruction rather than folding into the address the load reads
4758 // from. Two registers with no scale on either is the one addressing mode the rules have no
4759 // load through, because the folds that exist are the displacement one and the scaled ones,
4760 // and this is neither. That is a peephole worth having and not a thing this changes.
4761 assert_eq!(
4762 lower(&mut names, &source),
4763 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
4764 %1:gpr($rsi) = x64.arg_val_64\n %2:gpr(reuse 1) = x64.add_rr_64 %0, %1\n \
4765 %3:gpr = x64.mov_rm_32 [%2]\n x64.ret_val_32 %3($rax)\n}\n"
4766 );
4767 }
4768
4769 /// The address of a file scope name, which is what every use of a global and every string
4770 /// literal starts from.
4771 fn address_of(source: &mut Func, block: Block, names: &mut Interner, name: &str) -> Value {
4772 let symbol = names.intern(name);
4773 let mut build = Builder::new(source, block);
4774 build.value(
4775 InstData { extra: Extra::Symbol(symbol), ..InstData::new(Opcode::GlobalAddr) },
4776 Type::PTR,
4777 )
4778 }
4779
4780 #[test]
4781 fn the_address_of_a_name_is_one_instruction_carrying_the_name() {
4782 let (mut names, mut source, block, _) = blank(&[]);
4783 let counter = address_of(&mut source, block, &mut names, "counter");
4784 let mut build = Builder::new(&mut source, block);
4785 let loaded = build.load(Type::int(32), counter, plain(), Flags::default());
4786 build.ret(&[loaded]);
4787
4788 // `extern int counter; int f(void) { return counter; }`. The address is an addressing mode
4789 // that names no register and carries the symbol, which is what the assembler writes
4790 // relative to `%rip` and what the object writer leaves a relocation for.
4791 assert_eq!(
4792 lower(&mut names, &source),
4793 "mfunc @f {\nblock0:\n %0:gpr = x64.lea_64 [@counter]\n \
4794 %1:gpr = x64.mov_rm_32 [%0]\n x64.ret_val_32 %1($rax)\n}\n"
4795 );
4796 }
4797
4798 #[test]
4799 fn the_address_of_a_name_outside_the_file_is_read_out_of_the_offset_table() {
4800 let (mut names, mut source, block, _) = blank(&[]);
4801 let away = address_of(&mut source, block, &mut names, "away");
4802 Builder::new(&mut source, block).ret(&[away]);
4803 let elsewhere: Elsewhere = [names.intern("away")].into_iter().collect();
4804
4805 // `extern void away(void); void *f(void) { return away; }`. A load and not an address
4806 // computation, because the distance from here to a name a shared library may be the one
4807 // that defines is not a number any link can work out, and the slot the linker fills in is
4808 // in this program and so is a distance it has.
4809 let out =
4810 func(&source, &mut names, &SYSV, &elsewhere).expect("every instruction has a rule");
4811 assert_eq!(
4812 mir::print_func(&out.func, &names, ®S),
4813 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_rm_64 [got @away]\n \
4814 x64.ret_val_64 %0($rax)\n}\n"
4815 );
4816 }
4817
4818 #[test]
4819 fn the_address_of_a_thread_local_is_an_offset_out_of_the_table_plus_where_this_thread_starts() {
4820 let (mut names, mut source, block, _) = blank(&[]);
4821 let own = address_of(&mut source, block, &mut names, "own");
4822 Builder::new(&mut source, block).ret(&[own]);
4823 let elsewhere = Elsewhere::default().with_threads([names.intern("own")]);
4824
4825 // `extern _Thread_local int own; void *f(void) { return &own; }`. Three instructions where
4826 // the two cases above are one, because there is no address to load or to work out: the
4827 // slot holds how far into a thread's block the variable sits, `%fs:0` is where this
4828 // thread's block starts, and the sum of the two is this thread's copy.
4829 let out =
4830 func(&source, &mut names, &SYSV, &elsewhere).expect("every instruction has a rule");
4831 assert_eq!(
4832 mir::print_func(&out.func, &names, ®S),
4833 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_rm_64 [thread @own]\n \
4834 %1:gpr = x64.mov_rm_64 [fs:0]\n %2:gpr(reuse 1) = x64.add_rr_64 %0, %1\n \
4835 x64.ret_val_64 %2($rax)\n}\n"
4836 );
4837 }
4838
4839 /// The same load with nothing added to it, which is the whole of `__builtin_thread_pointer`.
4840 #[test]
4841 fn the_start_of_this_thread_s_own_storage_is_the_one_load_and_no_arithmetic() {
4842 let (mut names, mut source, block, _) = blank(&[]);
4843 let here =
4844 Builder::new(&mut source, block).value(InstData::new(Opcode::ThreadPointer), Type::PTR);
4845 Builder::new(&mut source, block).ret(&[here]);
4846
4847 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
4848 .expect("every instruction has a rule");
4849 assert_eq!(
4850 mir::print_func(&out.func, &names, ®S),
4851 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_rm_64 [fs:0]\n \
4852 x64.ret_val_64 %0($rax)\n}\n"
4853 );
4854 }
4855
4856 /// One `asm` statement, with its template and its constraint list written as a program does.
4857 fn assembly(
4858 source: &mut Func,
4859 block: Block,
4860 names: &mut Interner,
4861 template: &str,
4862 constraints: &str,
4863 args: &[Value],
4864 results: &[Type],
4865 ) -> Inst {
4866 clobbering(source, block, names, template, constraints, "memory", args, results)
4867 }
4868
4869 /// The same with a clobber list of its own, for the statements that are about one.
4870 #[allow(clippy::too_many_arguments)]
4871 fn clobbering(
4872 source: &mut Func,
4873 block: Block,
4874 names: &mut Interner,
4875 template: &str,
4876 constraints: &str,
4877 clobbers: &str,
4878 args: &[Value],
4879 results: &[Type],
4880 ) -> Inst {
4881 let info = AsmInfo {
4882 template: names.intern(template),
4883 constraints: names.intern(constraints),
4884 clobbers: names.intern(clobbers),
4885 targets: rucc_ir::BlockCallList::EMPTY,
4886 };
4887 Builder::new(source, block).inline_asm(info, args, results, Flags::VOLATILE)
4888 }
4889
4890 /// What a program asking the processor what it can do writes, which is the instruction whose
4891 /// every operand is a register its text does not name.
4892 #[test]
4893 fn a_template_whose_registers_are_named_by_the_constraints_places_them_from_the_letters() {
4894 let u32 = Type::int(32);
4895 let (mut names, mut source, block, _) = blank(&[]);
4896 let zero = Builder::new(&mut source, block).iconst(u32, 0);
4897 let out = clobbering(
4898 &mut source,
4899 block,
4900 &mut names,
4901 "cpuid",
4902 "=a,a",
4903 "ebx,ecx,edx",
4904 &[zero],
4905 &[u32],
4906 );
4907 let produced = source[out].results().next().expect("one result");
4908 Builder::new(&mut source, block).ret(&[produced]);
4909
4910 // `asm ("cpuid" : "=a" (n) : "a" (0) : "ebx", "ecx", "edx")`, which is the first thing
4911 // every program that has a faster path on some machines writes. Four registers written and
4912 // two read, none of them in the template, all of them out of the description, and the two
4913 // that the letters named are the statement's own. The subleaf is a zero because the
4914 // instruction reads `ecx` and the program said nothing about what is in it. The three
4915 // clobbers are gone because `cpuid` writes those three anyway, and saying it twice is one
4916 // register with two definitions.
4917 assert_eq!(
4918 lower(&mut names, &source),
4919 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_ri_32 0\n \
4920 %1:gpr = x64.mov_ri_64 0\n \
4921 %2:gpr($rax), %3:gpr($rbx), %4:gpr($rcx), %5:gpr($rdx) = x64.cpuid %0($rax), \
4922 %1($rcx)\n x64.ret_val_32 %2($rax)\n}\n"
4923 );
4924 }
4925
4926 /// A clobber the instruction does not write itself, which is the case the list is there for.
4927 /// It goes on as a definition of the register, in among the other definitions, because that is
4928 /// the whole of how a machine function says a register is not worth anything after this.
4929 #[test]
4930 fn a_clobber_the_instruction_does_not_write_itself_is_a_definition_of_that_register() {
4931 let (mut names, mut source, block, _) = blank(&[]);
4932 clobbering(&mut source, block, &mut names, "pause", "", "rsi,cc,memory", &[], &[]);
4933 Builder::new(&mut source, block).ret(&[]);
4934
4935 assert_eq!(lower(&mut names, &source), "mfunc @f {\nblock0:\n $rsi = x64.pause\n}\n");
4936 }
4937
4938 /// A clobber naming something this has no register for. Refused rather than dropped, since the
4939 /// list is the program saying which registers it may not leave anything in, and an entry
4940 /// nobody read is a register something may still be left in.
4941 #[test]
4942 fn a_clobber_this_has_no_register_for_is_refused() {
4943 let (mut names, mut source, block, _) = blank(&[]);
4944 clobbering(&mut source, block, &mut names, "pause", "", "zmm0", &[], &[]);
4945 Builder::new(&mut source, block).ret(&[]);
4946
4947 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
4948 .expect_err("there is no such register here");
4949 assert_eq!(
4950 failed.to_string(),
4951 "this `asm` says it destroys a register this has no name for"
4952 );
4953 }
4954
4955 #[test]
4956 fn an_asm_with_an_empty_template_and_no_operands_is_no_instructions() {
4957 let (mut names, mut source, block, _) = blank(&[]);
4958 assembly(&mut source, block, &mut names, "", "", &[], &[]);
4959 Builder::new(&mut source, block).ret(&[]);
4960
4961 // `asm volatile ("" : : : "memory")`, which is a barrier and nothing else. The barrier was
4962 // spent on the optimizer, which has finished by now, so what is left is nothing.
4963 assert_eq!(lower(&mut names, &source), "mfunc @f {\nblock0:\n}\n");
4964 }
4965
4966 #[test]
4967 fn an_output_an_input_is_tied_to_is_the_register_that_input_arrived_in() {
4968 let i32 = Type::int(32);
4969 let (mut names, mut source, block, args) = blank(&[i32]);
4970 let out = assembly(&mut source, block, &mut names, "", "=r,0", &args, &[i32]);
4971 let produced = source[out].results().next().expect("one result");
4972 Builder::new(&mut source, block).ret(&[produced]);
4973
4974 // `asm ("" : "=r" (x) : "0" (x))`, which is how a program stops the optimizer following a
4975 // value without changing it. The two share a place and the template writes nothing over
4976 // it, so the value comes back out of the register it went in.
4977 assert_eq!(
4978 lower(&mut names, &source),
4979 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
4980 x64.ret_val_32 %0($rax)\n}\n"
4981 );
4982 }
4983
4984 #[test]
4985 fn an_output_written_plus_is_the_same_rename() {
4986 let i32 = Type::int(32);
4987 let (mut names, mut source, block, args) = blank(&[i32]);
4988 let out = assembly(&mut source, block, &mut names, "", "+r", &args, &[i32]);
4989 let produced = source[out].results().next().expect("one result");
4990 Builder::new(&mut source, block).ret(&[produced]);
4991
4992 // `asm ("" : "+r" (x))`, which says the same thing in one operand instead of two.
4993 assert_eq!(
4994 lower(&mut names, &source),
4995 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_32\n \
4996 x64.ret_val_32 %0($rax)\n}\n"
4997 );
4998 }
4999
5000 #[test]
5001 fn an_output_nothing_is_tied_to_is_a_zero() {
5002 let i32 = Type::int(32);
5003 let (mut names, mut source, block, _) = blank(&[]);
5004 let out = assembly(&mut source, block, &mut names, "", "=r", &[], &[i32]);
5005 let produced = source[out].results().next().expect("one result");
5006 Builder::new(&mut source, block).ret(&[produced]);
5007
5008 // `asm ("" : "=r" (y))`, whose answer is whatever the assembly left in the register, and
5009 // an empty template leaves nothing. A definite value rather than a register nothing wrote,
5010 // because the allocator is owed a definition before the use however little the program is.
5011 assert_eq!(
5012 lower(&mut names, &source),
5013 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_ri_32 0\n x64.ret_val_32 %0($rax)\n}\n"
5014 );
5015 }
5016
5017 #[test]
5018 fn a_template_that_is_one_instruction_becomes_that_instruction() {
5019 let (mut names, mut source, block, _) = blank(&[]);
5020 assembly(&mut source, block, &mut names, "pause", "", &[], &[]);
5021 Builder::new(&mut source, block).ret(&[]);
5022
5023 // `asm volatile ("pause")`, which is what every spin lock in every allocator writes. One
5024 // instruction, no operands, and nothing between the template and the machine but the table
5025 // that already says what a `pause` is.
5026 assert_eq!(lower(&mut names, &source), "mfunc @f {\nblock0:\n x64.pause\n}\n");
5027 }
5028
5029 #[test]
5030 fn a_template_that_reads_a_segment_becomes_the_load_it_already_was() {
5031 let i64 = Type::int(64);
5032 let (mut names, mut source, block, _) = blank(&[]);
5033 let out = assembly(&mut source, block, &mut names, "movq %%fs:0, %0", "=r", &[], &[i64]);
5034 let produced = source[out].results().next().expect("one result");
5035 Builder::new(&mut source, block).ret(&[produced]);
5036
5037 // `asm ("movq %%fs:0, %0" : "=r" (tid))`, which is how a program finds the block its own
5038 // thread owns. The same instruction `crate::lower` already writes for a thread-local
5039 // variable, reached this time because a program wrote it out by hand.
5040 assert_eq!(
5041 lower(&mut names, &source),
5042 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_rm_64 [fs:0]\n \
5043 x64.ret_val_64 %0($rax)\n}\n"
5044 );
5045 }
5046
5047 #[test]
5048 fn a_template_naming_an_instruction_this_machine_has_not_got_is_refused() {
5049 let (mut names, mut source, block, _) = blank(&[]);
5050 assembly(&mut source, block, &mut names, "hcf", "", &[], &[]);
5051 Builder::new(&mut source, block).ret(&[]);
5052
5053 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
5054 .expect_err("there is no such instruction");
5055 assert_eq!(
5056 failed.to_string(),
5057 "this `asm` has instructions in its template, which nothing here assembles"
5058 );
5059 }
5060
5061 /// A register the template named is a claim on a register nobody told the allocator about.
5062 /// Refused rather than placed, because a register two things believe they own is a wrong
5063 /// program that nothing reports. A register a constraint letter names is a different thing and
5064 /// is placed, which the test above is about: there the statement said which of its own operands
5065 /// is in the register, and a name in the middle of a template says no such thing.
5066 #[test]
5067 fn a_template_naming_a_register_the_allocator_did_not_hand_out_is_refused() {
5068 let i64 = Type::int(64);
5069 let (mut names, mut source, block, _) = blank(&[]);
5070 let out = assembly(&mut source, block, &mut names, "movq %%rax, %0", "=r", &[], &[i64]);
5071 let produced = source[out].results().next().expect("one result");
5072 Builder::new(&mut source, block).ret(&[produced]);
5073
5074 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
5075 .expect_err("the template named a register");
5076 assert_eq!(failed.to_string(), "this `asm` has an operand this cannot place");
5077 }
5078
5079 #[test]
5080 fn a_constraint_list_that_does_not_describe_the_operands_is_refused() {
5081 let i32 = Type::int(32);
5082 let (mut names, mut source, block, args) = blank(&[i32]);
5083 assembly(&mut source, block, &mut names, "", "=r", &args, &[]);
5084 Builder::new(&mut source, block).ret(&[]);
5085
5086 // An output with no result to be, which is what the front end never writes and what a
5087 // hand written module can. Refused rather than placed by a guess.
5088 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
5089 .expect_err("the list and the instruction disagree");
5090 assert_eq!(failed.to_string(), "this `asm` has an operand this cannot place");
5091 }
5092
5093 /// A cast between a pointer and an integer, at whatever width the result is asked for.
5094 fn cast(source: &mut Func, block: Block, opcode: Opcode, from: Value, to: Type) -> Value {
5095 let mut build = Builder::new(source, block);
5096 let args = build.func().push_values(&[from]);
5097 build.value(InstData { args, ..InstData::new(opcode) }, to)
5098 }
5099
5100 #[test]
5101 fn a_cast_between_a_pointer_and_an_integer_as_wide_is_no_instruction_at_all() {
5102 let (mut names, mut source, block, args) = blank(&[Type::PTR]);
5103 let number = cast(&mut source, block, Opcode::PtrToInt, args[0], Type::int(64));
5104 Builder::new(&mut source, block).ret(&[number]);
5105
5106 // `long f(void *p) { return (long)p; }`. An address on this machine is an integer as wide
5107 // as the machine addresses, so the cast changes what the type system calls the value and
5108 // changes nothing about the value, and the register holding it is the one that held it.
5109 assert_eq!(
5110 lower(&mut names, &source),
5111 "mfunc @f {\nblock0:\n %0:gpr($rdi) = x64.arg_val_64\n \
5112 x64.ret_val_64 %0($rax)\n}\n"
5113 );
5114 }
5115
5116 #[test]
5117 fn a_null_pointer_is_a_constant_that_reaches_a_register_before_anything_reads_it() {
5118 let (mut names, mut source, block, _) = blank(&[]);
5119 let mut build = Builder::new(&mut source, block);
5120 let zero = build.iconst(Type::int(64), 0);
5121 let null = cast(&mut source, block, Opcode::IntToPtr, zero, Type::PTR);
5122 Builder::new(&mut source, block).ret(&[null]);
5123
5124 // `void *f(void) { return 0; }`. The cast is nothing, and reading its operand is what
5125 // writes the zero down: a constant is materialized where it is wanted rather than where
5126 // the IR defined it, and without the read there would be no instruction at all.
5127 assert_eq!(
5128 lower(&mut names, &source),
5129 "mfunc @f {\nblock0:\n %0:gpr = x64.mov_ri_64 0\n x64.ret_val_64 %0($rax)\n}\n"
5130 );
5131 }
5132
5133 #[test]
5134 fn the_five_linkages_the_ir_has_narrow_to_the_three_an_object_file_can_say() {
5135 let readings = [
5136 (Linkage::External, mir::Binding::Global),
5137 (Linkage::Common, mir::Binding::Global),
5138 (Linkage::Internal, mir::Binding::Local),
5139 (Linkage::Weak, mir::Binding::Weak),
5140 (Linkage::LinkOnce, mir::Binding::Weak),
5141 ];
5142 for (linkage, wanted) in readings {
5143 let (mut names, mut source, block, _) = blank(&[]);
5144 source.linkage = linkage;
5145 Builder::new(&mut source, block).ret(&[]);
5146 let out = func(&source, &mut names, &SYSV, &Elsewhere::default()).expect("a return");
5147 // The narrowing is done here rather than where the object is written, because a
5148 // machine function is all the assembler and the writer are ever handed.
5149 assert_eq!(out.func.binding, wanted, "{linkage:?}");
5150 }
5151 }
5152
5153 /// The visibility makes the same trip and is not narrowed on the way, because ELF says all
5154 /// three of them.
5155 ///
5156 /// Here for the reason the linkage above is here. A machine function is the whole of what the
5157 /// assembler and the object writer are handed, so a fact about the symbol that does not get
5158 /// onto one is a fact that is gone by the time anything could write it down, and the way that
5159 /// shows up is a shared library exporting the wrong set of names with nothing said anywhere.
5160 #[test]
5161 fn the_visibility_survives_the_trip_from_the_ir_to_a_machine_function() {
5162 let readings = [
5163 (Visibility::Default, mir::Visibility::Default),
5164 (Visibility::Hidden, mir::Visibility::Hidden),
5165 (Visibility::Protected, mir::Visibility::Protected),
5166 ];
5167 for (visibility, wanted) in readings {
5168 let (mut names, mut source, block, _) = blank(&[]);
5169 source.visibility = visibility;
5170 Builder::new(&mut source, block).ret(&[]);
5171 let out = func(&source, &mut names, &SYSV, &Elsewhere::default()).expect("a return");
5172 assert_eq!(out.func.visibility, wanted, "{visibility:?}");
5173 }
5174 }
5175
5176 #[test]
5177 fn a_cast_between_a_pointer_and_a_narrower_integer_is_reported() {
5178 let (mut names, mut source, block, args) = blank(&[Type::PTR]);
5179 let number = cast(&mut source, block, Opcode::PtrToInt, args[0], Type::int(32));
5180 Builder::new(&mut source, block).ret(&[number]);
5181
5182 // The front end never writes one: it casts at the address width and truncates or extends
5183 // around it, so both of those are the rules they always were. IR from somewhere else that
5184 // does write one is refused rather than compiled to a move that keeps the high half.
5185 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
5186 .expect_err("no rule narrows an address");
5187 assert_eq!(failed.to_string(), "no rule lowers a `ptrtoint` producing a `i32`");
5188 }
5189
5190 /// The type this machine has no register for.
5191 fn long_double() -> Type {
5192 Type::float(rucc_ir::Float::F80)
5193 }
5194
5195 #[test]
5196 fn a_double_widened_and_narrowed_again_goes_out_through_the_frame_and_back() {
5197 let f64 = Type::float(rucc_ir::Float::F64);
5198 let (mut names, mut source, block, args) = blank(&[f64]);
5199 let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
5200 let back = cast(&mut source, block, Opcode::FPTrunc, wide, f64);
5201 Builder::new(&mut source, block).ret(&[back]);
5202
5203 // `double f(double d) { long double x = d; return x; }`. The x87 reads memory and nothing
5204 // else, so the value is written to the crossing slot, loaded at the format that widens it
5205 // and put in the slot the eighty bit value lives in. Coming back is the same three the
5206 // other way. Both slots are addressed by a `lea` with nothing in it yet, which is what
5207 // every address in a frame looks like here until `finish` has the numbers.
5208 assert_eq!(
5209 lower(&mut names, &source),
5210 "mfunc @f {\nblock0:\n \
5211 %0:xmm($xmm0) = x64.arg_val_f64\n \
5212 %1:gpr = x64.lea_64 [$rsp]\n \
5213 %2:gpr = x64.lea_64 [$rsp]\n \
5214 x64.movsd_mr %0, [%1]\n \
5215 x64.fld_l [%1]\n \
5216 x64.fstp_t [%2]\n \
5217 %3:gpr = x64.lea_64 [$rsp]\n \
5218 %4:gpr = x64.lea_64 [$rsp]\n \
5219 x64.fld_t [%3]\n \
5220 x64.fstp_l [%4]\n \
5221 %5:xmm = x64.movsd_rm [%4]\n \
5222 x64.ret_val_f64 %5($xmm0)\n}\n"
5223 );
5224 }
5225
5226 #[test]
5227 fn a_long_double_has_sixteen_bytes_of_its_own_and_keeps_them() {
5228 let f64 = Type::float(rucc_ir::Float::F64);
5229 let (mut names, mut source, block, args) = blank(&[f64]);
5230 let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
5231 let once = cast(&mut source, block, Opcode::FPTrunc, wide, f64);
5232 let twice = cast(&mut source, block, Opcode::FPTrunc, wide, f64);
5233 let mut build = Builder::new(&mut source, block);
5234 let sum = build.binary(Opcode::FAdd, once, twice, Flags::default());
5235 build.ret(&[sum]);
5236
5237 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
5238 .expect("every instruction is written");
5239
5240 // Two slots and not four: sixteen bytes for the one eighty bit value, which is what the
5241 // psABI says one takes and is aligned to, and eight for the crossing, which every group
5242 // in the function shares because nothing is ever left in it. The value's slot is its own
5243 // for the whole function, so reading it twice reads the same sixteen bytes.
5244 assert_eq!(
5245 out.stack.locals,
5246 vec![Local { size: 8, align: 8 }, Local { size: 16, align: 16 }]
5247 );
5248 }
5249
5250 #[test]
5251 fn an_integer_becomes_a_long_double_by_being_loaded_as_one() {
5252 let (mut names, mut source, block, args) = blank(&[Type::int(64)]);
5253 let wide = cast(&mut source, block, Opcode::SIToFP, args[0], long_double());
5254 let back =
5255 cast(&mut source, block, Opcode::FPTrunc, wide, Type::float(rucc_ir::Float::F64));
5256 Builder::new(&mut source, block).ret(&[back]);
5257
5258 // `double f(long n) { long double x = n; return x; }`. `fild` is the same push at another
5259 // format, so the conversion is the load and there is no instruction that converts.
5260 let text = lower(&mut names, &source);
5261 assert!(text.contains("x64.mov_mr_64 %0, [%1]"), "{text}");
5262 assert!(text.contains("x64.fild_ll [%1]"), "{text}");
5263 }
5264
5265 #[test]
5266 fn a_long_double_becoming_an_integer_cuts_towards_zero_with_the_control_word() {
5267 let (mut names, mut source, block, args) = blank(&[Type::float(rucc_ir::Float::F64)]);
5268 let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
5269 let whole = cast(&mut source, block, Opcode::FPToSI, wide, Type::int(32));
5270 Builder::new(&mut source, block).ret(&[whole]);
5271
5272 // The one conversion here with no single instruction behind it. C cuts towards zero and
5273 // the unit rounds the way its control word says, so the word is saved, ORed with the two
5274 // bits that mean truncate, loaded, used and put back. Nine instructions for what `fisttp`
5275 // does in one, and `spec/10-backend.md` section 10.8 says why that one is not used.
5276 let text = lower(&mut names, &source);
5277 let group: Vec<&str> = text
5278 .lines()
5279 .map(str::trim)
5280 .filter(|line| line.starts_with("x64.f") || line.contains("_16"))
5281 .collect();
5282 assert_eq!(
5283 group,
5284 [
5285 "x64.fld_l [%1]",
5286 "x64.fstp_t [%2]",
5287 "x64.fnstcw [%5]",
5288 "%6:gpr = x64.mov_rm_16 [%5]",
5289 "%7:gpr(reuse 1) = x64.or_ri_16 %6, 3072",
5290 "x64.mov_mr_16 %7, [%5 + 2]",
5291 "x64.fldcw [%5 + 2]",
5292 "x64.fld_t [%3]",
5293 "x64.fistp_l [%4]",
5294 "x64.fldcw [%5]",
5295 ],
5296 "{text}"
5297 );
5298 }
5299
5300 #[test]
5301 fn a_long_double_is_read_and_written_as_the_bits_it_already_is() {
5302 let (mut names, mut source, block, args) = blank(&[Type::PTR, Type::PTR]);
5303 let mut build = Builder::new(&mut source, block);
5304 let value = build.load(long_double(), args[0], plain(), Flags::default());
5305 build.store(value, args[1], plain(), Flags::default());
5306 build.ret(&[]);
5307
5308 // `void f(long double *a, long double *b) { *b = *a; }`. A copy is a push and a pop at the
5309 // format the value is already in, which neither converts nor looks: a signalling NaN stays
5310 // one and nothing is raised, which is the whole of what makes it a copy.
5311 let text = lower(&mut names, &source);
5312 let group: Vec<&str> =
5313 text.lines().map(str::trim).filter(|line| line.starts_with("x64.f")).collect();
5314 assert_eq!(
5315 group,
5316 ["x64.fld_t [%0]", "x64.fstp_t [%2]", "x64.fld_t [%3]", "x64.fstp_t [%1]"],
5317 "{text}"
5318 );
5319 }
5320
5321 /// Two `long double` values, from two `double` parameters, and the instructions that made
5322 /// them, which every test below this one throws away.
5323 fn two_long_doubles(source: &mut Func, block: Block, args: &[Value]) -> (Value, Value) {
5324 let left = cast(source, block, Opcode::FPExt, args[0], long_double());
5325 let right = cast(source, block, Opcode::FPExt, args[1], long_double());
5326 (left, right)
5327 }
5328
5329 /// The x87 instructions of a function, in order, with everything else dropped.
5330 fn stack_only(text: &str) -> Vec<&str> {
5331 text.lines().map(str::trim).filter(|line| line.contains("x64.f")).collect()
5332 }
5333
5334 /// The two frame slots the last two addresses of a function were taken of, which in a
5335 /// comparison are the two operands in the order they go on the stack.
5336 fn pushed(out: &Lowered) -> Vec<usize> {
5337 let taken: Vec<usize> = out.stack.addresses.iter().map(|&(_, local)| local).collect();
5338 taken[taken.len() - 2..].to_vec()
5339 }
5340
5341 #[test]
5342 fn adding_two_long_doubles_pushes_both_and_leaves_the_answer_in_a_slot() {
5343 let f64 = Type::float(rucc_ir::Float::F64);
5344 let (mut names, mut source, block, args) = blank(&[f64, f64]);
5345 let (left, right) = two_long_doubles(&mut source, block, &args);
5346 let sum =
5347 Builder::new(&mut source, block).binary(Opcode::FAdd, left, right, Flags::default());
5348 let back = cast(&mut source, block, Opcode::FPTrunc, sum, f64);
5349 Builder::new(&mut source, block).ret(&[back]);
5350
5351 // `double f(double a, double b) { return (long double) a + (long double) b; }`. The last
5352 // four lines are the add: both operands pushed, the instruction that names neither of
5353 // them because they are the top two of a stack, and the answer taken off into its slot.
5354 let text = lower(&mut names, &source);
5355 assert_eq!(
5356 stack_only(&text),
5357 [
5358 "x64.fld_l [%2]",
5359 "x64.fstp_t [%3]",
5360 "x64.fld_l [%4]",
5361 "x64.fstp_t [%5]",
5362 "x64.fld_t [%6]",
5363 "x64.fld_t [%7]",
5364 "x64.fadd_p",
5365 "x64.fstp_t [%8]",
5366 "x64.fld_t [%9]",
5367 "x64.fstp_l [%10]",
5368 ],
5369 "{text}"
5370 );
5371 }
5372
5373 #[test]
5374 fn a_subtraction_pushes_the_left_operand_first_and_asks_for_the_att_spelling() {
5375 let f64 = Type::float(rucc_ir::Float::F64);
5376 let (mut names, mut source, block, args) = blank(&[f64, f64]);
5377 let (left, right) = two_long_doubles(&mut source, block, &args);
5378 let less =
5379 Builder::new(&mut source, block).binary(Opcode::FSub, left, right, Flags::default());
5380 let back = cast(&mut source, block, Opcode::FPTrunc, less, f64);
5381 Builder::new(&mut source, block).ret(&[back]);
5382
5383 // The left one goes on first, so it ends up under the right one, and the answer wanted is
5384 // the one below minus the top. In AT&T that is `fsubrp`, since `fsubp` there is `DE E0+i`
5385 // and computes the other one. The `r` says which spelling this is and not which order the
5386 // pushes were in. `crates/rucc/tests/x87.rs` is what says the answer is right, because a
5387 // name is what got this wrong the first time.
5388 let text = lower(&mut names, &source);
5389 assert_eq!(
5390 &stack_only(&text)[4..8],
5391 ["x64.fld_t [%6]", "x64.fld_t [%7]", "x64.fsubr_p", "x64.fstp_t [%8]"],
5392 "{text}"
5393 );
5394 }
5395
5396 #[test]
5397 fn negating_a_long_double_turns_the_sign_over_and_reads_nothing() {
5398 let f64 = Type::float(rucc_ir::Float::F64);
5399 let (mut names, mut source, block, args) = blank(&[f64]);
5400 let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
5401 let flipped = Builder::new(&mut source, block).unary(Opcode::FNeg, wide, long_double());
5402 let back = cast(&mut source, block, Opcode::FPTrunc, flipped, f64);
5403 Builder::new(&mut source, block).ret(&[back]);
5404
5405 // `fchs` and not a subtraction from zero, which would give a different answer at a negative
5406 // zero and would signal at a NaN. It does not read the value as a number at all.
5407 let text = lower(&mut names, &source);
5408 assert_eq!(
5409 &stack_only(&text)[2..5],
5410 ["x64.fld_t [%3]", "x64.fchs", "x64.fstp_t [%4]"],
5411 "{text}"
5412 );
5413 }
5414
5415 #[test]
5416 fn comparing_two_long_doubles_puts_the_left_one_on_top() {
5417 let f64 = Type::float(rucc_ir::Float::F64);
5418 let (mut names, mut source, block, args) = blank(&[f64, f64]);
5419 let (left, right) = two_long_doubles(&mut source, block, &args);
5420 let mut build = Builder::new(&mut source, block);
5421 build.fcmp(FloatPred::Ogt, left, right, Flags::default());
5422 build.ret(&[]);
5423
5424 // `a > b`. `fucomip` asks about the top of the stack against what is under it, so the
5425 // operand the predicate is about has to go on last, which is the other way round from the
5426 // arithmetic above. The pop that clears the loser and the byte that reads the flags are
5427 // both inside the one opcode.
5428 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
5429 .expect("every instruction is written");
5430 let slots = pushed(&out);
5431 assert_eq!(slots, [2, 1], "the right operand goes on first and the left one on top");
5432 let text = mir::print_func(&out.func, &names, ®S);
5433 assert_eq!(
5434 &stack_only(&text)[4..],
5435 ["x64.fld_t [%6]", "x64.fld_t [%7]", "%8:gpr = x64.fucomip_set_a"],
5436 "{text}"
5437 );
5438 }
5439
5440 #[test]
5441 fn a_comparison_that_the_machine_has_backwards_swaps_the_two_pushes() {
5442 let f64 = Type::float(rucc_ir::Float::F64);
5443 let (mut names, mut source, block, args) = blank(&[f64, f64]);
5444 let (left, right) = two_long_doubles(&mut source, block, &args);
5445 let mut build = Builder::new(&mut source, block);
5446 build.fcmp(FloatPred::Olt, left, right, Flags::default());
5447 build.ret(&[]);
5448
5449 // `a < b` is `b > a` and this machine has the one condition, so the same opcode runs with
5450 // the operands the other way round. The same trade the vector rules make, and it has to
5451 // be the same one: a `long double` comparison that picked a different condition from the
5452 // `double` comparison of the same two numbers would be wrong at exactly the unordered
5453 // cases the two conditions differ on.
5454 //
5455 // Which slot each push names is the whole of the difference from the test above, and the
5456 // text does not show it, since an address in a frame is a `lea` with nothing in it until
5457 // `finish` has the numbers. So the slots are what is read here.
5458 let out = func(&source, &mut names, &SYSV, &Elsewhere::default())
5459 .expect("every instruction is written");
5460 let slots = pushed(&out);
5461 assert_eq!(slots, [1, 2], "the left operand goes on first and the right one on top");
5462 let text = mir::print_func(&out.func, &names, ®S);
5463 assert_eq!(
5464 &stack_only(&text)[4..],
5465 ["x64.fld_t [%6]", "x64.fld_t [%7]", "%8:gpr = x64.fucomip_set_a"],
5466 "{text}"
5467 );
5468 }
5469
5470 #[test]
5471 fn an_ordered_equal_needs_a_second_byte_to_put_the_two_conditions_together() {
5472 let f64 = Type::float(rucc_ir::Float::F64);
5473 let (mut names, mut source, block, args) = blank(&[f64, f64]);
5474 let (left, right) = two_long_doubles(&mut source, block, &args);
5475 let mut build = Builder::new(&mut source, block);
5476 build.fcmp(FloatPred::Oeq, left, right, Flags::default());
5477 build.ret(&[]);
5478
5479 // Equal and ordered are two conditions and the flags carry both, so the opcode writes a
5480 // second register as well as the one the value is in and ANDs them together. Said here by
5481 // handing it a spare, since an instruction that wrote a register nothing knew about would
5482 // be an instruction the allocator could put a live value in the way of.
5483 let text = lower(&mut names, &source);
5484 assert!(text.contains("%8:gpr, %9:gpr = x64.fucomip_set_e_and_np"), "{text}");
5485 }
5486
5487 #[test]
5488 fn a_comparison_that_is_never_asked_is_reported() {
5489 let f64 = Type::float(rucc_ir::Float::F64);
5490 let (mut names, mut source, block, args) = blank(&[f64, f64]);
5491 let (left, right) = two_long_doubles(&mut source, block, &args);
5492 let mut build = Builder::new(&mut source, block);
5493 build.fcmp(FloatPred::False, left, right, Flags::default());
5494 build.ret(&[]);
5495
5496 // Always false is a constant and not a comparison, so there is no condition to pick and
5497 // nothing here folds it into one: an instruction that quietly agreed with it would hide
5498 // that the optimizer left a comparison in that it should have taken out.
5499 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
5500 .expect_err("no condition is always false");
5501 assert_eq!(failed.to_string(), "no rule lowers a `fcmp` producing a `i1`");
5502 }
5503
5504 #[test]
5505 fn a_long_double_constant_is_the_bits_of_it_put_where_the_value_lives() {
5506 let (mut names, mut source, block, args) = blank(&[Type::PTR]);
5507 let mut build = Builder::new(&mut source, block);
5508 // `1.5L`, which is the leading bit and one more of significand, and an exponent of zero.
5509 let one_and_a_half = build.fconst(long_double(), 0x3fff_c000_0000_0000_0000);
5510 build.store(one_and_a_half, args[0], plain(), Flags::default());
5511 build.ret(&[]);
5512
5513 // No x87 instruction at all. A slot holding one of these is the value, so a constant is
5514 // its ten bytes written where the value lives, and whatever reads it does the `fld`.
5515 let text = lower(&mut names, &source);
5516 assert!(text.contains("x64.mov_ri_64 -4611686018427387904"), "{text}");
5517 assert!(text.contains("x64.mov_ri_16 16383"), "{text}");
5518 assert!(text.contains("x64.mov_mr_16 %3, [%1 + 8]"), "{text}");
5519 // The six bytes above the ten are the padding that makes the type sixteen wide, and they
5520 // are unspecified rather than zero, so nothing writes them.
5521 assert_eq!(text.matches("x64.mov_mr").count(), 2, "{text}");
5522 }
5523
5524 #[test]
5525 fn a_negative_long_double_constant_keeps_the_bit_above_its_exponent() {
5526 let (mut names, mut source, block, args) = blank(&[Type::PTR]);
5527 let mut build = Builder::new(&mut source, block);
5528 let minus = build.fconst(long_double(), 0xbfff_c000_0000_0000_0000);
5529 build.store(minus, args[0], plain(), Flags::default());
5530 build.ret(&[]);
5531
5532 // `-1.5L`. The sign is the top bit of the two byte half, so the immediate that half is put
5533 // in a register with is above the signed range of sixteen bits and has to stay there: read
5534 // as a number it would be negative, and it is not a number, it is two bytes.
5535 let text = lower(&mut names, &source);
5536 assert!(text.contains("x64.mov_ri_16 49151"), "{text}");
5537 }
5538
5539 #[test]
5540 fn a_long_double_crosses_an_edge_as_an_address_and_is_copied_where_it_lands() {
5541 let (mut names, mut source, block, args) = blank(&[Type::float(rucc_ir::Float::F64)]);
5542 let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
5543 let next = source.create_block();
5544 let param = source.append_param(next, long_double());
5545 Builder::new(&mut source, block).jump(next, &[wide]);
5546 Builder::new(&mut source, next).ret(&[param]);
5547
5548 // What the edge carries is the address of the slot the value is already in, which is an
5549 // ordinary register the allocator has an opinion about. The block on the other side copies
5550 // the sixteen bytes into a slot of its own before anything reads them, so a second edge
5551 // handing over a second address would still leave one place for a reader to look.
5552 let text = lower(&mut names, &source);
5553 let second: Vec<&str> = text
5554 .lines()
5555 .skip_while(|line| !line.starts_with("block1"))
5556 .skip(1)
5557 .take(3)
5558 .map(str::trim)
5559 .collect();
5560 assert_eq!(
5561 second,
5562 ["x64.fld_t [%4]", "%5:gpr = x64.lea_64 [$rsp]", "x64.fstp_t [%5]"],
5563 "{text}"
5564 );
5565 }
5566
5567 #[test]
5568 fn more_long_doubles_at_a_block_than_the_stack_is_deep_are_reported() {
5569 let f64 = Type::float(rucc_ir::Float::F64);
5570 let (mut names, mut source, block, args) = blank(&[f64]);
5571 let wide = cast(&mut source, block, Opcode::FPExt, args[0], long_double());
5572 let next = source.create_block();
5573 let params: Vec<Value> =
5574 (0..=X87_DEPTH).map(|_| source.append_param(next, long_double())).collect();
5575 let carried: Vec<Value> = params.iter().map(|_| wide).collect();
5576 Builder::new(&mut source, block).jump(next, &carried);
5577 Builder::new(&mut source, next).ret(&[params[0]]);
5578
5579 // The copies go through the x87 stack so that every one of them is read before any of them
5580 // is written, which is what makes a block that swaps two of these right. Nine of them do
5581 // not fit on the stack, and copying the ninth before or after the rest is the order that
5582 // could be wrong, so it is refused instead.
5583 let failed = func(&source, &mut names, &SYSV, &Elsewhere::default())
5584 .expect_err("nine do not fit on the stack");
5585 assert_eq!(
5586 failed.to_string(),
5587 "block1 takes 9 parameters of type `f80` and only 8 can cross an edge at once"
5588 );
5589 assert_eq!(failed.inst(), None);
5590 }
5591}