A 2D array

Memory is a line of bytes. It has no idea what a row is, and there is no arrangement of hardware that will give it one. A grid of three rows and four columns is therefore twelve numbers in a row, plus an agreement about how to read them.

The agreement here is row major: the whole of row 0, then the whole of row 1, then row 2. Under that agreement the element at row r and column c is number r * COLS + c in the line, and the multiplication is where the shape of the grid actually lives.

r9 is 26, which is 5 + 6 + 7 + 8, a whole row. r10 is 21, which is 3 + 7 + 11, a whole column. The two loops that produced them are the same length and do very different amounts of work.

The row loop adds 1 to its index each pass, so it reads twelve consecutive qwords going forwards. The column loop adds COLS to its index, so it jumps thirty two bytes at a time. On real hardware that is several times slower for exactly the same arithmetic, because neighbouring elements of a row arrive in the same cache line and are already there by the time the loop asks for them, while neighbouring elements of a column are each in a different one. A program that walks a large array the wrong way round can spend most of its time waiting for memory.

Nothing recorded that this was a grid, which you can prove. Change COLS equ 4 to COLS equ 3 and do not touch a byte of the data. The program now reads the same twelve numbers as four rows of three, r8 becomes 8, and no error is reported anywhere, because the only thing that ever said "four columns" was the imul.

The three dq lines are one array for the same reason: the assembler writes twelve qwords one after another and the line breaks in the source exist for your benefit only.

[grid + rax*8] scales by 8 because an element is a qword. An element of some other size, a twenty byte record say, needs a real multiplication into a register first, because the scale in an address only goes up to 8.