===== Instruction Set =====
Alpha is a 64-bit load/store RISC architecture. It has 32 integer and 32 floating-point registers, all 64 bits wide, and every instruction is 32 bits long. Memory is accessed only by loads and stores, which in the original architecture work on 32-bit longwords and 64-bit quadwords only. Later processors added extensions for byte and word access, multimedia, floating-point moves and square roots, and bit counting. This page summarizes the registers and their conventional use, the extensions and the processors that introduced them, and the instructions and assembler features most often met in Linux programming. The {{wiki:documentation:references:alpha_architecture_reference_manual_4th_edition.pdf?linkonly|Alpha Architecture Reference Manual, Fourth Edition}} is the definitive reference; see [[documentation:references|References]].
==== Registers ====
=== Integer ===
There are 32 integer registers. The ''$31'' register holds the constant value 0. Writes to it are ignored.
By convention 12 registers are for temporary values, 7 (''$9'' to ''$15'') are saved across function calls, and 6 are for passing arguments to functions. The return address (''$26'') and the stack pointer (''$30'') are also preserved.
^ Register ^ Alternate Name ^ Preserved? ^ Purpose ^
| **$0** | **$v0** | No | Return Value |
| **$1**–**$8** | **$t0**–**$t7** | No | Temporary Registers |
| **$9**–**$14** | **$s0**–**$s5** | Yes | Saved Registers |
| **$15** | **$s6** or **$fp** | Yes | Saved Register or Frame Pointer |
| **$16**–**$21** | **$a0**–**$a5** | No | Function Arguments |
| **$22**–**$25** | **$t8**–**$t11** | No | Temporary Registers |
| **$26** | **$ra** | Yes | Function Return Address |
| **$27** | **$pv** or **$t12** | No | Procedure Value or Temporary Register |
| **$28** | **$at** | No | Reserved for Assembler |
| **$29** | **$gp** | No | Global Pointer |
| **$30** | **$sp** | Yes | Stack Pointer |
| **$31** | **$zero** | | Zero Sink |
=== Floating-Point ===
There are 32 floating-point registers. The ''$f31'' register holds the constant value 0.0. Writes to it are ignored.
By convention, 15 registers are for temporary values, 8 are saved across function calls, and 6 are for passing arguments to functions.
^ Register ^ Preserved? ^ Purpose ^
| **$f0** | No | Return Value |
| **$f1** | No | Return Value of Imaginary Part |
| **$f2**–**$f9** | Yes | Saved Registers |
| **$f10**–**$f15** | No | Temporary Registers |
| **$f16**–**$f21** | No | Function Arguments |
| **$f22**–**$f30** | No | Temporary Registers |
| **$f31** | | Zero Sink |
==== Example program ====
The program below calls ''printf'' from assembly. It is built and run with ''gcc hello.S -o hello'' and ''./hello''. The ''.frame'', ''.mask'', and ''.prologue'' directives describe the stack frame for debuggers and the unwinder, and ''ldgp'' sets up the global pointer, which must be reloaded after every call to another function. [(>[[https://sourceware.org/git/?p=binutils-gdb.git;a=blob;f=gas/doc/c-alpha.texi|gas/doc/c-alpha.texi]], binutils)]
.data
PRINT: .asciz "Hello, World!\n"
.text
.align 4
.set noreorder
.arch ev56
.globl main
.ent main
main:
.frame $sp,16,$26,0 # 16-byte frame, return address in $26
.mask 0x4000000,-16 # $26 is saved in the frame
ldgp $gp,0($27) # load global pointer
lda $sp,-16($sp) # allocate a stack frame
stq $26,0($sp) # save return address to stack
.prologue 1 # end of prologue; 1 means it uses $gp
lda $16,PRINT # load format string
# $16 is first argument to functions
jsr $26,printf # call printf
ldgp $gp,0($26) # reload global pointer
# (necessary after function calls)
mov $31,$0 # return val = 0
ldq $26,0($sp) # load return address from stack
lda $sp,16($sp) # release the stack frame
ret $31,($26),1 # return, (1 signifies return from a procedure)
.end main
.section .note.GNU-stack,"",@progbits
==== ISA extensions ====
The extensions accumulate along the line EV56 (21164A, BWX), PCA56 (21164PC, adds MVI), EV6 (21264, adds FIX), and EV67 (21264A, adds CIX); each processor in that sequence implements the extensions of those before it, so the 21164A has BWX but not MVI. The [[amask]] instruction reports at run time which extensions a processor implements, using the bit given in the table; ''IMPLVER'' reports only the processor generation, for tuning decisions. [(>{{wiki:documentation:references:alpha_architecture_reference_manual_4th_edition.pdf?linkonly|Alpha Architecture Reference Manual, Fourth Edition}}, section 4.11 and Appendix D, pp. 4-141, D-4 to D-5)] GCC enables an extension with its ''-m'' option, or with a ''-mcpu'' value for a processor that has it. [(>[[https://gcc.gnu.org/git/?p=gcc.git;a=blob;f=gcc/config/alpha/alpha.cc|gcc/config/alpha/alpha.cc]], GCC)] [(>[[https://gcc.gnu.org/git/?p=gcc.git;a=blob;f=gcc/config/alpha/alpha.opt|gcc/config/alpha/alpha.opt]], GCC)]
^ Extension ^ First processor ^ AMASK bit ^ GCC option ^ First ''-mcpu'' ^
| Byte/word extension (BWX) | EV56 (21164A) | 0 | ''-mbwx'' | ''ev56'' |
| Motion video instructions (MVI) | PCA56 (21164PC) | 8 | ''-mmax'' | ''pca56'' |
| Floating-point extension (FIX) | EV6 (21264) | 1 | ''-mfix'' | ''ev6'' |
| Count extension (CIX) | EV67 (21264A) | 2 | ''-mcix'' | ''ev67'' |
Linux requires BWX since Linux 6.10; see [[documentation:kernel#processor_support|Kernel: Processor support]].
=== Byte/word extension (BWX) ===
The [[wp>DEC_Alpha#Byte-Word_Extensions_(BWX)|byte/word extension]] adds loads and stores of 8-bit bytes and 16-bit words, without which they are done by loading the containing quadword and extracting or inserting the byte; see [[documentation:porting:byte_word_access|Byte and Word Access]].
^ Mnemonic ^ Description ^
| **ldbu** | Load byte, zero-extended |
| **ldwu** | Load word, zero-extended |
| **sextb** | Sign-extend byte |
| **sextw** | Sign-extend word |
| **stb** | Store byte |
| **stw** | Store word |
=== Motion video instructions (MVI) ===
The [[wp>DEC_Alpha#Motion_Video_Instructions_(MVI)|motion video instructions]] operate on bytes and words packed into a 64-bit register.
^ Mnemonic ^ Description ^
| **maxsb8** | Vector signed byte maximum |
| **maxsw4** | Vector signed word maximum |
| **maxub8** | Vector unsigned byte maximum |
| **maxuw4** | Vector unsigned word maximum |
| **minsb8** | Vector signed byte minimum |
| **minsw4** | Vector signed word minimum |
| **minub8** | Vector unsigned byte minimum |
| **minuw4** | Vector unsigned word minimum |
| **perr** | Pixel error |
| **pklb** | Pack longwords to bytes |
| **pkwb** | Pack words to bytes |
| **unpkbl** | Unpack bytes to longwords |
| **unpkbw** | Unpack bytes to words |
=== Floating-point extension (FIX) ===
The [[wp>DEC_Alpha#Floating-point_Extensions_(FIX)|floating-point extension]] adds moves between the integer and floating-point registers, which otherwise go through memory, and square root instructions.
^ Mnemonic ^ Description ^ Format ^
| **itofs** | Copy the low 32 bits of an integer register to a floating-point register in S_floating format | IEEE |
| **itoft** | Copy 64 bits from an integer register to a floating-point register | IEEE |
| **itoff** | Copy the low 32 bits of an integer register to a floating-point register in F_floating format | VAX |
| **ftois** | Copy a floating-point register in S_floating format to an integer register | IEEE |
| **ftoit** | Copy 64 bits from a floating-point register to an integer register | IEEE |
| **sqrts** | Square root of an S_floating value | IEEE |
| **sqrtt** | Square root of a T_floating value | IEEE |
| **sqrtf** | Square root of an F_floating value | VAX |
| **sqrtg** | Square root of a G_floating value | VAX |
=== Count extension (CIX) ===
The [[wp>DEC_Alpha#Count_Extensions_(CIX)|count extension]] adds bit counting instructions.
^ Mnemonic ^ Description ^
| **ctlz** | Count leading zeros |
| **ctpop** | Count population (the number of 1 bits) |
| **cttz** | Count trailing zeros |
==== Atomic operations ====
Alpha has no atomic read-modify-write instructions. Atomic operations are built from a load-locked and store-conditional pair: ''ldl_l'' or ''ldq_l'' loads a longword or quadword and sets a lock flag, and ''stl_c'' or ''stq_c'' stores only if no other processor has written to the locked location in the meantime, writing 1 to its source register on success and 0 on failure. The sequence is retried in a loop until the store succeeds. There are no byte or word forms, even with BWX. The architecture limits what may appear between the two instructions, and the ordering of memory accesses around them is set with ''mb'' and ''wmb'' barriers. [(>{{wiki:documentation:references:alpha_architecture_reference_manual_4th_edition.pdf?linkonly|Alpha Architecture Reference Manual, Fourth Edition}}, sections 4.2.4 and 4.2.5, pp. 4-9 to 4-15)] See [[documentation:porting:memory_model|Memory Model]] for the rules and their consequences for Linux code.
==== Integer division ====
Alpha has no integer divide instruction. [(>[[https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/arch/alpha/lib/divide.S?h=v7.3-rc1|arch/alpha/lib/divide.S]], Linux 7.3)] GCC compiles integer ''/'' and ''%'' with a variable divisor into calls to helper routines: ''%%__divq%%'', ''%%__divqu%%'', ''%%__remq%%'' and ''%%__remqu%%'' for 64-bit operands, and ''%%__divl%%'', ''%%__divlu%%'', ''%%__reml%%'' and ''%%__remlu%%'' for 32-bit ones. [(>[[https://gcc.gnu.org/git/?p=gcc.git;a=blob;f=gcc/config/alpha/alpha.md|gcc/config/alpha/alpha.md]], GCC)] In user space the C library provides them; they are exported from ''libc.so.6.1''. [(>[[https://sourceware.org/git/?p=glibc.git;a=blob;f=sysdeps/alpha/Versions|sysdeps/alpha/Versions]], glibc)] The kernel has its own copies.
The helpers are not ordinary functions. The dividend is passed in ''$24'' (''t10'') and the divisor in ''$25'' (''t11''), the result is returned in ''$27'' (''t12''), the return address is in ''$23'' (''t9''), and only ''$27'' and ''$28'' (''at'') may be changed; every other register is preserved. [(>[[https://sourceware.org/git/?p=glibc.git;a=blob;f=sysdeps/alpha/div_libc.h|sysdeps/alpha/div_libc.h]], glibc)] A JIT compiler or hand-written assembly that needs division has to call them with this convention or divide by other means. glibc's versions use the floating-point divider where the operands allow an exact result, and fall back to a shift-and-subtract loop otherwise. [(>[[https://sourceware.org/git/?p=glibc.git;a=blob;f=sysdeps/alpha/divq.S|sysdeps/alpha/divq.S]], glibc)]
Division by zero is detected in software: the helper executes the ''gentrap'' PALcode call with the code ''GEN_INTDIV'', and the kernel delivers ''SIGFPE'' with ''si_code'' ''FPE_INTDIV''. [(>[[https://sourceware.org/git/?p=glibc.git;a=blob;f=sysdeps/alpha/div_libc.h|sysdeps/alpha/div_libc.h]], glibc)] [(>[[https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/arch/alpha/kernel/traps.c?h=v7.3-rc1|arch/alpha/kernel/traps.c]], Linux 7.3)]
==== PALcode calls ====
''call_pal'' calls a function of the PALcode, the firmware layer beneath the operating system. Linux uses the PALcode interface of Tru64 UNIX (OSF/1). The calls a user program meets are: [(>[[https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/arch/alpha/include/uapi/asm/pal.h?h=v7.3-rc1|arch/alpha/include/uapi/asm/pal.h]], Linux 7.3)]
^ Call ^ Number ^ Use ^
| ''callsys'' | 0x83 | Enter the kernel for a system call |
| ''bpt'' | 0x80 | Breakpoint trap, used by debuggers |
| ''imb'' | 0x86 | Instruction memory barrier, needed after writing code to memory, as a JIT compiler does |
| ''rduniq'' | 0x9e | Read the thread's unique value, the thread pointer used for TLS |
| ''wruniq'' | 0x9f | Write the thread's unique value |
GCC's ''%%__builtin_thread_pointer%%'' compiles to ''call_pal 0x9e'' (''rduniq''), and glibc uses it to find the thread's TLS block. [(>[[https://gcc.gnu.org/git/?p=gcc.git;a=blob;f=gcc/config/alpha/alpha.md|gcc/config/alpha/alpha.md]], GCC)] [(>[[https://sourceware.org/git/?p=glibc.git;a=blob;f=sysdeps/alpha/nptl/tls.h|sysdeps/alpha/nptl/tls.h]], glibc)] The system call convention, with the number in ''$0'', arguments in ''$16'' to ''$21'', and an error flag returned in ''$19'', is described under [[documentation:porting:abi#system_calls|Linux ABI Differences: System calls]].
==== Relocation operators ====
GNU as accepts explicit relocation annotations on individual instructions, written after the operands with ''!''. GCC emits them when it schedules the address calculations itself, and they appear in its assembly output and in hand-written assembly: [(>[[https://sourceware.org/git/?p=binutils-gdb.git;a=blob;f=gas/doc/c-alpha.texi|gas/doc/c-alpha.texi]], binutils)]
* ''!literal'' marks an ''ldq'' that loads a symbol's address from the global offset table (GOT), and the ''!lituse_*'' operators mark the instructions that use the loaded address, so that the linker can optimize the sequence.
* ''!gpdisp'' marks the ''ldah'' and ''lda'' pair that computes the global pointer, as the ''ldgp'' macro does.
* ''!gprelhigh'', ''!gprellow'', and ''!gprel'' address data at a fixed offset from the global pointer.
* ''!tlsgd'', ''!tlsldm'', ''!gotdtprel'', ''!dtprelhi'', ''!gottprel'', ''!tprelhi'' and related operators implement the thread-local storage models.
ldah $29,0($27) !gpdisp!1
lda $29,0($29) !gpdisp!1
ldq $1,b($29) !literal!2
ldl $2,0($1) !lituse_base!2
''!literal'' and ''!gprel'' reach the GOT or data through a 16-bit displacement from the global pointer, which limits them to a 64 KiB window around the global pointer; see [[documentation:toolchains|Toolchains]] for the "relocation truncated to fit" errors this causes in large programs.
==== Further reading ====
Raymond Chen's blog series about the Alpha is a detailed, Windows-focused discussion of the instruction set.
* [[https://devblogs.microsoft.com/oldnewthing/20170807-00/?p=96766|The Alpha AXP, part 1: Initial plunge]]
* [[https://devblogs.microsoft.com/oldnewthing/20170808-00/?p=96775|The Alpha AXP, part 2: Integer calculations]]
* [[https://devblogs.microsoft.com/oldnewthing/20170809-00/?p=96785|The Alpha AXP, part 3: Integer constants]]
* [[https://devblogs.microsoft.com/oldnewthing/20170810-00/?p=96795|The Alpha AXP, part 4: Bit 15. Ugh. Bit 15.]]
* [[https://devblogs.microsoft.com/oldnewthing/20170811-00/?p=96805|The Alpha AXP, part 5: Conditional operations and control flow]]
* [[https://devblogs.microsoft.com/oldnewthing/20170814-00/?p=96806|The Alpha AXP, part 6: Memory access, basics]]
* [[https://devblogs.microsoft.com/oldnewthing/20170815-00/?p=96816|The Alpha AXP, part 7: Memory access, loading unaligned data]]
* [[https://devblogs.microsoft.com/oldnewthing/20170816-00/?p=96825|The Alpha AXP, part 8: Memory access, storing bytes and words and unaligned data]]
* [[https://devblogs.microsoft.com/oldnewthing/20170817-00/?p=96835|The Alpha AXP, part 9: The memory model and atomic memory operations]]
* [[https://devblogs.microsoft.com/oldnewthing/20170818-00/?p=96845|The Alpha AXP, part 10: Atomic updates to byte and word memory units]]
* [[https://devblogs.microsoft.com/oldnewthing/20170821-00/?p=96855|The Alpha AXP, part 11: Processor faults]]
* [[https://devblogs.microsoft.com/oldnewthing/20170822-00/?p=96865|The Alpha AXP, part 12: How you detect carry on a processor with no carry?]]
* [[https://devblogs.microsoft.com/oldnewthing/20170823-00/?p=96875|The Alpha AXP, part 13: On treating a 64-bit processor as if it were a 32-bit processor]]
* [[https://devblogs.microsoft.com/oldnewthing/20170825-00/?p=96887|The Alpha AXP, part 14: On the strange behavior of writes to the zero register]]
* [[https://devblogs.microsoft.com/oldnewthing/20170828-00/?p=96895|The Alpha AXP, part 15: Variadic functions]]
* [[https://devblogs.microsoft.com/oldnewthing/20170829-00/?p=96897|The Alpha AXP, part 16: What are the dire consequences of having 32-bit values in non-canonical form?]]
* [[https://devblogs.microsoft.com/oldnewthing/20170830-00/?p=96906|The Alpha AXP, part 17: Reconstructing a call stack]]
{{tag>documentation}}