Alpha is a 64-bit load/store RISC architecture. It has 32 integer and 32 floating-point registers, all 64 bits wide, and every instruction is 32 bits long. Memory is accessed only by loads and stores, which in the original architecture work on 32-bit longwords and 64-bit quadwords only. Later processors added extensions for byte and word access, multimedia, floating-point moves and square roots, and bit counting. This page summarizes the registers and their conventional use, the extensions and the processors that introduced them, and the instructions and assembler features most often met in Linux programming. The Alpha Architecture Reference Manual, Fourth Edition is the definitive reference; see References.
There are 32 integer registers. The $31 register holds the constant value 0. Writes to it are ignored.
By convention 12 registers are for temporary values, 7 ($9 to $15) are saved across function calls, and 6 are for passing arguments to functions. The return address ($26) and the stack pointer ($30) are also preserved.
| Register | Alternate Name | Preserved? | Purpose |
|---|---|---|---|
| $0 | $v0 | No | Return Value |
| $1–$8 | $t0–$t7 | No | Temporary Registers |
| $9–$14 | $s0–$s5 | Yes | Saved Registers |
| $15 | $s6 or $fp | Yes | Saved Register or Frame Pointer |
| $16–$21 | $a0–$a5 | No | Function Arguments |
| $22–$25 | $t8–$t11 | No | Temporary Registers |
| $26 | $ra | Yes | Function Return Address |
| $27 | $pv or $t12 | No | Procedure Value or Temporary Register |
| $28 | $at | No | Reserved for Assembler |
| $29 | $gp | No | Global Pointer |
| $30 | $sp | Yes | Stack Pointer |
| $31 | $zero | Zero Sink |
There are 32 floating-point registers. The $f31 register holds the constant value 0.0. Writes to it are ignored.
By convention, 15 registers are for temporary values, 8 are saved across function calls, and 6 are for passing arguments to functions.
| Register | Preserved? | Purpose |
|---|---|---|
| $f0 | No | Return Value |
| $f1 | No | Return Value of Imaginary Part |
| $f2–$f9 | Yes | Saved Registers |
| $f10–$f15 | No | Temporary Registers |
| $f16–$f21 | No | Function Arguments |
| $f22–$f30 | No | Temporary Registers |
| $f31 | Zero Sink |
The program below calls printf from assembly. It is built and run with gcc hello.S -o hello and ./hello. The .frame, .mask, and .prologue directives describe the stack frame for debuggers and the unwinder, and ldgp sets up the global pointer, which must be reloaded after every call to another function. 1)
.data PRINT: .asciz "Hello, World!\n" .text .align 4 .set noreorder .arch ev56 .globl main .ent main main: .frame $sp,16,$26,0 # 16-byte frame, return address in $26 .mask 0x4000000,-16 # $26 is saved in the frame ldgp $gp,0($27) # load global pointer lda $sp,-16($sp) # allocate a stack frame stq $26,0($sp) # save return address to stack .prologue 1 # end of prologue; 1 means it uses $gp lda $16,PRINT # load format string # $16 is first argument to functions jsr $26,printf # call printf ldgp $gp,0($26) # reload global pointer # (necessary after function calls) mov $31,$0 # return val = 0 ldq $26,0($sp) # load return address from stack lda $sp,16($sp) # release the stack frame ret $31,($26),1 # return, (1 signifies return from a procedure) .end main .section .note.GNU-stack,"",@progbits
The extensions accumulate along the line EV56 (21164A, BWX), PCA56 (21164PC, adds MVI), EV6 (21264, adds FIX), and EV67 (21264A, adds CIX); each processor in that sequence implements the extensions of those before it, so the 21164A has BWX but not MVI. The amask instruction reports at run time which extensions a processor implements, using the bit given in the table; IMPLVER reports only the processor generation, for tuning decisions. 2) GCC enables an extension with its -m option, or with a -mcpu value for a processor that has it. 3) 4)
| Extension | First processor | AMASK bit | GCC option | First -mcpu |
|---|---|---|---|---|
| Byte/word extension (BWX) | EV56 (21164A) | 0 | -mbwx | ev56 |
| Motion video instructions (MVI) | PCA56 (21164PC) | 8 | -mmax | pca56 |
| Floating-point extension (FIX) | EV6 (21264) | 1 | -mfix | ev6 |
| Count extension (CIX) | EV67 (21264A) | 2 | -mcix | ev67 |
Linux requires BWX since Linux 6.10; see Kernel: Processor support.
The byte/word extension adds loads and stores of 8-bit bytes and 16-bit words, without which they are done by loading the containing quadword and extracting or inserting the byte; see Byte and Word Access.
| Mnemonic | Description |
|---|---|
| ldbu | Load byte, zero-extended |
| ldwu | Load word, zero-extended |
| sextb | Sign-extend byte |
| sextw | Sign-extend word |
| stb | Store byte |
| stw | Store word |
The motion video instructions operate on bytes and words packed into a 64-bit register.
| Mnemonic | Description |
|---|---|
| maxsb8 | Vector signed byte maximum |
| maxsw4 | Vector signed word maximum |
| maxub8 | Vector unsigned byte maximum |
| maxuw4 | Vector unsigned word maximum |
| minsb8 | Vector signed byte minimum |
| minsw4 | Vector signed word minimum |
| minub8 | Vector unsigned byte minimum |
| minuw4 | Vector unsigned word minimum |
| perr | Pixel error |
| pklb | Pack longwords to bytes |
| pkwb | Pack words to bytes |
| unpkbl | Unpack bytes to longwords |
| unpkbw | Unpack bytes to words |
The floating-point extension adds moves between the integer and floating-point registers, which otherwise go through memory, and square root instructions.
| Mnemonic | Description | Format |
|---|---|---|
| itofs | Copy the low 32 bits of an integer register to a floating-point register in S_floating format | IEEE |
| itoft | Copy 64 bits from an integer register to a floating-point register | IEEE |
| itoff | Copy the low 32 bits of an integer register to a floating-point register in F_floating format | VAX |
| ftois | Copy a floating-point register in S_floating format to an integer register | IEEE |
| ftoit | Copy 64 bits from a floating-point register to an integer register | IEEE |
| sqrts | Square root of an S_floating value | IEEE |
| sqrtt | Square root of a T_floating value | IEEE |
| sqrtf | Square root of an F_floating value | VAX |
| sqrtg | Square root of a G_floating value | VAX |
The count extension adds bit counting instructions.
| Mnemonic | Description |
|---|---|
| ctlz | Count leading zeros |
| ctpop | Count population (the number of 1 bits) |
| cttz | Count trailing zeros |
Alpha has no atomic read-modify-write instructions. Atomic operations are built from a load-locked and store-conditional pair: ldl_l or ldq_l loads a longword or quadword and sets a lock flag, and stl_c or stq_c stores only if no other processor has written to the locked location in the meantime, writing 1 to its source register on success and 0 on failure. The sequence is retried in a loop until the store succeeds. There are no byte or word forms, even with BWX. The architecture limits what may appear between the two instructions, and the ordering of memory accesses around them is set with mb and wmb barriers. 5) See Memory Model for the rules and their consequences for Linux code.
Alpha has no integer divide instruction. 6) GCC compiles integer / and % with a variable divisor into calls to helper routines: __divq, __divqu, __remq and __remqu for 64-bit operands, and __divl, __divlu, __reml and __remlu for 32-bit ones. 7) In user space the C library provides them; they are exported from libc.so.6.1. 8) The kernel has its own copies.
The helpers are not ordinary functions. The dividend is passed in $24 (t10) and the divisor in $25 (t11), the result is returned in $27 (t12), the return address is in $23 (t9), and only $27 and $28 (at) may be changed; every other register is preserved. 9) A JIT compiler or hand-written assembly that needs division has to call them with this convention or divide by other means. glibc's versions use the floating-point divider where the operands allow an exact result, and fall back to a shift-and-subtract loop otherwise. 10)
Division by zero is detected in software: the helper executes the gentrap PALcode call with the code GEN_INTDIV, and the kernel delivers SIGFPE with si_code FPE_INTDIV. 11) 12)
call_pal calls a function of the PALcode, the firmware layer beneath the operating system. Linux uses the PALcode interface of Tru64 UNIX (OSF/1). The calls a user program meets are: 13)
| Call | Number | Use |
|---|---|---|
callsys | 0x83 | Enter the kernel for a system call |
bpt | 0x80 | Breakpoint trap, used by debuggers |
imb | 0x86 | Instruction memory barrier, needed after writing code to memory, as a JIT compiler does |
rduniq | 0x9e | Read the thread's unique value, the thread pointer used for TLS |
wruniq | 0x9f | Write the thread's unique value |
GCC's __builtin_thread_pointer compiles to call_pal 0x9e (rduniq), and glibc uses it to find the thread's TLS block. 14) 15) The system call convention, with the number in $0, arguments in $16 to $21, and an error flag returned in $19, is described under Linux ABI Differences: System calls.
GNU as accepts explicit relocation annotations on individual instructions, written after the operands with !. GCC emits them when it schedules the address calculations itself, and they appear in its assembly output and in hand-written assembly: 16)
!literal marks an ldq that loads a symbol's address from the global offset table (GOT), and the !lituse_* operators mark the instructions that use the loaded address, so that the linker can optimize the sequence.!gpdisp marks the ldah and lda pair that computes the global pointer, as the ldgp macro does.!gprelhigh, !gprellow, and !gprel address data at a fixed offset from the global pointer.!tlsgd, !tlsldm, !gotdtprel, !dtprelhi, !gottprel, !tprelhi and related operators implement the thread-local storage models.ldah $29,0($27) !gpdisp!1 lda $29,0($29) !gpdisp!1 ldq $1,b($29) !literal!2 ldl $2,0($1) !lituse_base!2
!literal and !gprel reach the GOT or data through a 16-bit displacement from the global pointer, which limits them to a 64 KiB window around the global pointer; see Toolchains for the "relocation truncated to fit" errors this causes in large programs.
Raymond Chen's blog series about the Alpha is a detailed, Windows-focused discussion of the instruction set.