User Tools

Site Tools


documentation:instruction_set

Instruction Set

Alpha is a 64-bit load/store RISC architecture. It has 32 integer and 32 floating-point registers, all 64 bits wide, and every instruction is 32 bits long. Memory is accessed only by loads and stores, which in the original architecture work on 32-bit longwords and 64-bit quadwords only. Later processors added extensions for byte and word access, multimedia, floating-point moves and square roots, and bit counting. This page summarizes the registers and their conventional use, the extensions and the processors that introduced them, and the instructions and assembler features most often met in Linux programming. The Alpha Architecture Reference Manual, Fourth Edition is the definitive reference; see References.

Registers

Integer

There are 32 integer registers. The $31 register holds the constant value 0. Writes to it are ignored.

By convention 12 registers are for temporary values, 7 ($9 to $15) are saved across function calls, and 6 are for passing arguments to functions. The return address ($26) and the stack pointer ($30) are also preserved.

Register Alternate Name Preserved? Purpose
$0 $v0 No Return Value
$1–$8 $t0–$t7 No Temporary Registers
$9–$14 $s0–$s5 Yes Saved Registers
$15 $s6 or $fp Yes Saved Register or Frame Pointer
$16–$21 $a0–$a5 No Function Arguments
$22–$25 $t8–$t11 No Temporary Registers
$26 $ra Yes Function Return Address
$27 $pv or $t12 No Procedure Value or Temporary Register
$28 $at No Reserved for Assembler
$29 $gp No Global Pointer
$30 $sp Yes Stack Pointer
$31 $zero Zero Sink

Floating-Point

There are 32 floating-point registers. The $f31 register holds the constant value 0.0. Writes to it are ignored.

By convention, 15 registers are for temporary values, 8 are saved across function calls, and 6 are for passing arguments to functions.

Register Preserved? Purpose
$f0 No Return Value
$f1 No Return Value of Imaginary Part
$f2–$f9 Yes Saved Registers
$f10–$f15 No Temporary Registers
$f16–$f21 No Function Arguments
$f22–$f30 No Temporary Registers
$f31 Zero Sink

Example program

The program below calls printf from assembly. It is built and run with gcc hello.S -o hello and ./hello. The .frame, .mask, and .prologue directives describe the stack frame for debuggers and the unwinder, and ldgp sets up the global pointer, which must be reloaded after every call to another function. 1)

	.data
PRINT:	.asciz	"Hello, World!\n"
 
	.text
	.align	4
	.set	noreorder
	.arch	ev56
	.globl	main
	.ent	main
main:
	.frame	$sp,16,$26,0		# 16-byte frame, return address in $26
	.mask	0x4000000,-16		# $26 is saved in the frame
	ldgp	$gp,0($27)		# load global pointer
	lda	$sp,-16($sp)		# allocate a stack frame
	stq	$26,0($sp)		# save return address to stack
	.prologue 1			# end of prologue; 1 means it uses $gp
 
	lda	$16,PRINT		# load format string
					# $16 is first argument to functions
	jsr	$26,printf		# call printf
	ldgp	$gp,0($26)		# reload global pointer
					# (necessary after function calls)
 
	mov	$31,$0			# return val = 0
	ldq	$26,0($sp)		# load return address from stack
	lda	$sp,16($sp)		# release the stack frame
	ret	$31,($26),1		# return, (1 signifies return from a procedure)
	.end	main
 
	.section .note.GNU-stack,"",@progbits

ISA extensions

The extensions accumulate along the line EV56 (21164A, BWX), PCA56 (21164PC, adds MVI), EV6 (21264, adds FIX), and EV67 (21264A, adds CIX); each processor in that sequence implements the extensions of those before it, so the 21164A has BWX but not MVI. The amask instruction reports at run time which extensions a processor implements, using the bit given in the table; IMPLVER reports only the processor generation, for tuning decisions. 2) GCC enables an extension with its -m option, or with a -mcpu value for a processor that has it. 3) 4)

Extension First processor AMASK bit GCC option First -mcpu
Byte/word extension (BWX) EV56 (21164A) 0 -mbwx ev56
Motion video instructions (MVI) PCA56 (21164PC) 8 -mmax pca56
Floating-point extension (FIX) EV6 (21264) 1 -mfix ev6
Count extension (CIX) EV67 (21264A) 2 -mcix ev67

Linux requires BWX since Linux 6.10; see Kernel: Processor support.

Byte/word extension (BWX)

The byte/word extension adds loads and stores of 8-bit bytes and 16-bit words, without which they are done by loading the containing quadword and extracting or inserting the byte; see Byte and Word Access.

Mnemonic Description
ldbu Load byte, zero-extended
ldwu Load word, zero-extended
sextb Sign-extend byte
sextw Sign-extend word
stb Store byte
stw Store word

Motion video instructions (MVI)

The motion video instructions operate on bytes and words packed into a 64-bit register.

Mnemonic Description
maxsb8 Vector signed byte maximum
maxsw4 Vector signed word maximum
maxub8 Vector unsigned byte maximum
maxuw4 Vector unsigned word maximum
minsb8 Vector signed byte minimum
minsw4 Vector signed word minimum
minub8 Vector unsigned byte minimum
minuw4 Vector unsigned word minimum
perr Pixel error
pklb Pack longwords to bytes
pkwb Pack words to bytes
unpkbl Unpack bytes to longwords
unpkbw Unpack bytes to words

Floating-point extension (FIX)

The floating-point extension adds moves between the integer and floating-point registers, which otherwise go through memory, and square root instructions.

Mnemonic Description Format
itofs Copy the low 32 bits of an integer register to a floating-point register in S_floating format IEEE
itoft Copy 64 bits from an integer register to a floating-point register IEEE
itoff Copy the low 32 bits of an integer register to a floating-point register in F_floating format VAX
ftois Copy a floating-point register in S_floating format to an integer register IEEE
ftoit Copy 64 bits from a floating-point register to an integer register IEEE
sqrts Square root of an S_floating value IEEE
sqrtt Square root of a T_floating value IEEE
sqrtf Square root of an F_floating value VAX
sqrtg Square root of a G_floating value VAX

Count extension (CIX)

The count extension adds bit counting instructions.

Mnemonic Description
ctlz Count leading zeros
ctpop Count population (the number of 1 bits)
cttz Count trailing zeros

Atomic operations

Alpha has no atomic read-modify-write instructions. Atomic operations are built from a load-locked and store-conditional pair: ldl_l or ldq_l loads a longword or quadword and sets a lock flag, and stl_c or stq_c stores only if no other processor has written to the locked location in the meantime, writing 1 to its source register on success and 0 on failure. The sequence is retried in a loop until the store succeeds. There are no byte or word forms, even with BWX. The architecture limits what may appear between the two instructions, and the ordering of memory accesses around them is set with mb and wmb barriers. 5) See Memory Model for the rules and their consequences for Linux code.

Integer division

Alpha has no integer divide instruction. 6) GCC compiles integer / and % with a variable divisor into calls to helper routines: __divq, __divqu, __remq and __remqu for 64-bit operands, and __divl, __divlu, __reml and __remlu for 32-bit ones. 7) In user space the C library provides them; they are exported from libc.so.6.1. 8) The kernel has its own copies.

The helpers are not ordinary functions. The dividend is passed in $24 (t10) and the divisor in $25 (t11), the result is returned in $27 (t12), the return address is in $23 (t9), and only $27 and $28 (at) may be changed; every other register is preserved. 9) A JIT compiler or hand-written assembly that needs division has to call them with this convention or divide by other means. glibc's versions use the floating-point divider where the operands allow an exact result, and fall back to a shift-and-subtract loop otherwise. 10)

Division by zero is detected in software: the helper executes the gentrap PALcode call with the code GEN_INTDIV, and the kernel delivers SIGFPE with si_code FPE_INTDIV. 11) 12)

PALcode calls

call_pal calls a function of the PALcode, the firmware layer beneath the operating system. Linux uses the PALcode interface of Tru64 UNIX (OSF/1). The calls a user program meets are: 13)

Call Number Use
callsys 0x83 Enter the kernel for a system call
bpt 0x80 Breakpoint trap, used by debuggers
imb 0x86 Instruction memory barrier, needed after writing code to memory, as a JIT compiler does
rduniq 0x9e Read the thread's unique value, the thread pointer used for TLS
wruniq 0x9f Write the thread's unique value

GCC's __builtin_thread_pointer compiles to call_pal 0x9e (rduniq), and glibc uses it to find the thread's TLS block. 14) 15) The system call convention, with the number in $0, arguments in $16 to $21, and an error flag returned in $19, is described under Linux ABI Differences: System calls.

Relocation operators

GNU as accepts explicit relocation annotations on individual instructions, written after the operands with !. GCC emits them when it schedules the address calculations itself, and they appear in its assembly output and in hand-written assembly: 16)

  • !literal marks an ldq that loads a symbol's address from the global offset table (GOT), and the !lituse_* operators mark the instructions that use the loaded address, so that the linker can optimize the sequence.
  • !gpdisp marks the ldah and lda pair that computes the global pointer, as the ldgp macro does.
  • !gprelhigh, !gprellow, and !gprel address data at a fixed offset from the global pointer.
  • !tlsgd, !tlsldm, !gotdtprel, !dtprelhi, !gottprel, !tprelhi and related operators implement the thread-local storage models.
	ldah	$29,0($27)	!gpdisp!1
	lda	$29,0($29)	!gpdisp!1
	ldq	$1,b($29)	!literal!2
	ldl	$2,0($1)	!lituse_base!2

!literal and !gprel reach the GOT or data through a 16-bit displacement from the global pointer, which limits them to a 64 KiB window around the global pointer; see Toolchains for the "relocation truncated to fit" errors this causes in large programs.

Further reading

documentation/instruction_set.txt · Last modified: by 127.0.0.1