diff --git a/20161001.html b/20161001.html new file mode 100644 index 0000000..7c8e1d7 --- /dev/null +++ b/20161001.html @@ -0,0 +1,660 @@ +
==============
+
+ FPGA NOTES
+
+==============
+
+===============
+ CORE BUDGET
+===============
+ 1 BRAM (32-bit x 1024 entry, with 2 ports both which can read or write)
+ 2 DSPs
+400 LUTs (50 slices, 8 LUTs per slice)
+
+
+==============
+ LUTs / MUX
+==============
+1 LUT = 2:1 MUX x2 (get two of these)
+1 LUT = 4:1 MUX
+2 LUT = 8:1 MUX
+4 LUT = 16:1 MUX
+8 LUT = 32:1 MUX
+
+
+===================================
+ LUTs / GENERAL PURPOSE FUNCTION
+===================================
+1 LUT = 5:1 x2 (get two of these sharing same 5 input bits)
+1 LUT = 6:1
+2 LUT = 13:1
+4 LUT = 27:1
+
+
+===============
+ SLICE RULES
+===============
+Carry chain for slice can start at LUT 0 or LUT 4
+Distributed RAM granularity is half a slice (starting at LUT 0 or LUT 4)
+
+
+========================
+ DSP PATTERN DETECTOR
+========================
+From docs, "use of the pattern detector leads to a moderate speed reduction
+ due to the extra logic on the pattern detect path"
+Suggests not using this to check for zero
+ So branch on signed or unsigned only?
+Wouldn't be able to use this for both saturation checks and zero check anyway
+
+
+====================
+
+ FUNCTIONAL UNITS
+
+====================
+
+============================
+ CURRENT LUT BUDGET USAGE
+============================
+LUTs % Usage
+==== === =====
+28 7 Program counter
+12 3 Return stack
+107 27 BRAM variable bit-width windows (includes adr^imm)
+64 16 Register file
+---- --- -----
+211 53 Total
+
+
+===================
+ PROGRAM COUNTER
+===================
+10-bit program counter (PC)
+Only lower 8-bits of PC increment on linear execution
+ Requires only an 8-bit PC+1 computation (one slice)
+
+Aim to minimize critical path getting next address to BRAM
+ Only one level of LUT to compute next address
+ All inputs registered at end of prior clock
+ Followed by 8-bit add to compute possible PC for next clock
+
+PC function inputs per output bit (map to 13:1 function at 2 LUTs/bit)
+ Bits Meaning
+ ==== =======
+ 1 Next PC if not branching (computed in prior clock)
+ 1 Top of return stack
+ 1 Immediate absolute branch address
+ 1 PSP P output register (computed branch target in prior clock)
+ 1 PSP P output register sign bit (for conditional branch)
+ 8 Up to 8 bits from instruction opcode to decode
+
+LUTs Usage
+==== =====
+20 13:1 function for next 10-bit PC computation including instruction decode
+8 PC+1 adder for 8 lower bits of PC
+---- -----
+28 Total (7% of 400 LUT/BRAM budget)
+
+
+================
+ RETURN STACK
+================
+Going to plan on a dedicated return stack for now
+Only need single port for return stack (either call/push, or return/pop)
+Hardware background,
+ Distributed RAM works in 4 LUT granularity
+ 1 LUT provides 2x SPRAM32 (single port 32x1 RAM)
+ Writes are synchronous on clock edge
+ Reads are async
+ 8 LUTs for a 32 x 16-bit return stack data
+
+Not using everything
+ Padding to 4 LUT granularity
+ Only need 10-bits out of 16-bits (6-bits free for other state)
+ Likely going to keep only 4-bit top of stack address register
+ Could use other 16 entries for run function on new message?
+
+Todo
+ Control inputs, adder input, etc
+
+LUTs Usage
+==== =====
+8 32 entry x 16-bit return stack
+4 4-bit top of stack pointer
+?
+---- -----
+12 Total (3% of 400 LUT/BRAM budget)
+
+
+=================================================
+ ADDRESS REGISTER XOR INSTEAD OF ADD IMMEDIATE
+=================================================
+Planning on [address ^ immediate] addressing
+ This removes an adder from the design
+
+XOR is the same as [address + immediate] for an n-bit immediate
+ When lower n-bits of address are zero
+ Means data must be aligned to the nearest pow2 of maximum immediate offset
+
+Using the following terms,
+ ggggggoooo
+ g = group bits (address bits choose the group of data)
+ o = offset bits (address bits are zero, immediate chooses element in group)
+
+For bits in address which are not cleared (ie the group bits),
+ Setting bits in the immediate results in accessing a neighbor group
+ Regardless of the starting group in the address register
+ It is possible to roll through all aligned groups
+ But ordering is different based on starting group address
+ Example of group bits for address crossed with immediate
+ 00 01 10 11
+ +-------------
+ 00 | 00 01 10 11
+ 01 | 01 00 11 10
+ 10 | 10 11 00 01
+ 11 | 11 10 01 00
+
+
+===================================
+ BRAM VARIABLE BIT-WIDTH WINDOWS
+===================================
+Trying to support transparent pack/unpack of variable bit-widths from BRAM
+ Want zero impact to ISA, no special instructions
+ Instead dividing address range into windows of different bit-widths
+ Each address range addresses at a multiple of the bit-width
+ Effectively the high bits of address choose the bit-width
+
+Store path limited to {8,16,32}-bit
+ Only using BRAM byte write mask to avoid any {read, modify, write}
+
+Fixed signed vs unsigned configuration
+ 32-bit doesn't matter
+ 16-bit going to go with signed (needed for vector or audio)
+ 8-bit unsigned (not so sure about that)
+ 4-bit unsigned for sure (sprites?)
+
+BRAMs always in 32-bit port mode,
+ fedcba9876543210
+ ================
+ .xxxxxxxxxx00000 - requires 10-bit address
+
+Address register,
+ fedcba9876543210
+ ================
+ 00....xxxxxxxxxx - 1024 x 32-bit
+ 01...xxxxxxxxxxx - 2048 x 16-bit
+ 10..xxxxxxxxxxxx - 4096 x 8-bit
+ 11.xxxxxxxxxxxxx - 8192 x 4-bit (supported for read only)
+
+Address register value to BRAM address translation
+ This needs to include XORing the immediate
+ Uses a 13:1 function for each bit,
+ Bits Meaning
+ ==== =======
+ 4 Address shifted left {0,1,2,3} bits
+ 4 Immediate shifted left {0,1,2,3} bits
+ 2 The 'fe' address bits
+ 3 Up to 3 bits from instruction opcode to decode
+ Should select if use immediate, etc
+
+Permutations (showing address and byte write mask for store),
+ 11111111111111110000000000000000
+ fedcba9876543210fedcba9876543210
+ ================================
+ aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa - 32-bit adr=00....xxxxxxxxxx write=1111
+ ................bbbbbbbbbbbbbbbb - 16-bit adr=01...xxxxxxxxxx0 write=0011
+ cccccccccccccccc................ - 16-bit adr=01...xxxxxxxxxx1 write=1100
+ ........................dddddddd - 8-bit adr=10..xxxxxxxxxx00 write=0001
+ ................eeeeeeee........ - 8-bit adr=10..xxxxxxxxxx01 write=0010
+ ........ffffffff................ - 8-bit adr=10..xxxxxxxxxx10 write=0100
+ gggggggg........................ - 8-bit adr=10..xxxxxxxxxx11 write=1000
+ ............................hhhh - 4-bit adr=11.xxxxxxxxxx000
+ ........................iiii.... - 4-bit adr=11.xxxxxxxxxx001
+ ....................jjjj........ - 4-bit adr=11.xxxxxxxxxx010
+ ................kkkk............ - 4-bit adr=11.xxxxxxxxxx011
+ ............llll................ - 4-bit adr=11.xxxxxxxxxx100
+ ........mmmm.................... - 4-bit adr=11.xxxxxxxxxx101
+ ....nnnn........................ - 4-bit adr=11.xxxxxxxxxx110
+ oooo............................ - 4-bit adr=11.xxxxxxxxxx111
+
+Shift value for store,
+ Requires 3:1 MUX per bit, 32 LUTs
+
+Generate write enable for store,
+ Requires same 4-bits per function,
+ 2 lower address bits
+ 2 upper address bits
+ 2 LUTs (5:1 function sharing inputs, 2 outputs per LUT)
+
+Unpack after load,
+ 11111111111111110000000000000000
+ fedcba9876543210fedcba9876543210
+ ================================
+ aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa
+ <---------------bbbbbbbbbbbbbbbb - sign extended
+ <---------------cccccccccccccccc - sign extended
+ 000000000000000000000000dddddddd
+ 000000000000000000000000eeeeeeee
+ 000000000000000000000000ffffffff
+ 000000000000000000000000gggggggg
+ 0000000000000000000000000000hhhh
+ 0000000000000000000000000000iiii
+ 0000000000000000000000000000jjjj
+ 0000000000000000000000000000kkkk
+ 0000000000000000000000000000llll
+ 0000000000000000000000000000mmmm
+ 0000000000000000000000000000nnnn
+ 0000000000000000000000000000oooo
+ ================================
+ xxxxxxxxxxxxxxxx................ - 4:1 MUX/bit (a bit, top b bit, top c bit, 0)
+ ................xxxxxxxx........ - 4:1 MUX/bit ({a,b,c} bit, 0)
+ ........................xxxx.... - 8:1 MUX/bit ({a,b,...g} bit, 0)
+ ............................xxxx - 15:1 MUX/bit ({a,b,...o} bit)
+
+Unpack if done with MUX control logic computed in prior clock,
+ This registers 4-bits extra, to reduce LUT cost on MUX
+ Register 2-bits for top 24-bit MUX control
+ Logic function (3:1 sharing inputs), one LUT
+ up lo 1 0
+ == === = =
+ 00 xxx | 0 0
+ 01 xx0 | 0 1
+ 01 xx1 | 1 0
+ 10 xxx | 1 1
+ 11 xxx | 1 1
+ Register 3-bits for 2nd lowest 4-bit MUX control
+ Logic function (4:1 sharing inputs), rounds to 2 LUTs
+ up lo 2 1 0
+ == == = = =
+ 00 xxx | 0 0 0
+ 01 xx0 | 0 0 1
+ 01 xx1 | 0 1 0
+ 10 x00 | 0 1 1
+ 10 x01 | 1 0 0
+ 10 x10 | 1 0 1
+ 10 x11 | 1 1 0
+ 11 xxx | 1 1 1
+ Register 4-bits for lower 4-bit MUX control
+ Logic function (5:1 sharing inputs), 2 LUTs
+ Skipping as pattern is obvious ...
+
+LUTs Usage
+==== =====
+20 Address register value to BRAM address translation 2 LUTs x 10-bits
+32 Shift value for store
+2 Generate write enable
+5 Unpack MUX control logic (done on clock computing address for BRAM)
+24 Unpack MUX for 24-bits
+8 Unpack MUX for 4-bits (higher nibble)
+16 Unpack MUX for 4-bits (lower nibble)
+---- -----
+107 Total (27% of 400 LUT/BRAM budget)
+
+
+=================
+ REGISTER FILE
+=================
+Using the smallest possible, but assuming the need for 1 write and 3 read ports,
+ 8 LUTs (one SLICEM) yields one Quad-port 32 entry x 4-bit RAM
+
+Register file write sources,
+ DSP P output register
+ BRAM load
+ What else?
+ TODO?
+ Need to fold these choices into BRAM unpack logic?
+ Or is BRAM load also forwarded into DSP inputs without first going to reg file?
+
+Register file read sources,
+ Address register
+ DSP A,B,C inputs
+ What else?
+
+LUTs Usage
+==== =====
+64 Total (16% of 400 LUT/BRAM budget)
+
+
+=======
+ DSP
+=======
+Not attempting to use all features of DSP
+ Pre-add 'a+d' tossed because of extra pipeline stage
+
+DSP is effectively modal,
+ If multiply is enabled then there are only 2 options,
+ 'p = c+(a*b)'
+ 'p = c-(a*b)' <- is multiply subtract worth it? (assuming yes for now)
+ Otherwise,
+ 'p = c OP (a:b)
+
+
+===================
+ MESSAGE PASSING
+===================
+Todo
+
+
+=================================
+
+ ISA / CODE GENERATION DETAILS
+
+=================================
+Todo
+
+=======
+ ISA
+=======
+Going to get messy for now.
+Attempting first to describe everything which needs instruction control.
+This will overflow a 32-bit instruction.
+In process culling options which are less needed.
+In hope of eventually fitting everything.
+
+Source input data which can be accessed each clock,
+ Register file loads (from prior clock),
+ 32-bit x
+ 32-bit y
+ 32-bit z
+ Instruction immediate for this instruction,
+ 10-bit i
+ BRAM last unpacked fetch (from prior clock),
+ 32-bit f
+ DSP output (from prior clock),
+ 32|48-bit p
+
+Sink output data,
+ DSP inputs,
+ 18-bit b
+ 24-bit a (could be up to 30-bit, but no LUTs for that)
+ 32-bit c (could be up to 48-bit, but no LUTs for that)
+ ?-bit control bits (todo)
+ BRAM output word
+ 32-bit o
+ BRAM memory address before translation
+ 16-bit m
+ BRAM access control bits
+ Register file entries for each port
+ 5-bit immediate s (store port)
+ 5-bit immediate t
+ 5-bit immediate u
+ 5-bit immediate v
+ Register file write value
+ 32-bit w
+
+Alphabet usage,
+ abcdefghijklmnopqrstuvwxyz
+ ==========================
+ abc....................... DSP inputs
+ ...............p.......... DSP output
+ ........i................. immediate
+ ............m............. BRAM memory address
+ .....f.................... BRAM fetched value
+ ..............o........... BRAM output word
+ ..................s....... Register file store port
+ ...................tuv.... Register file read ports
+ ......................w... Register file write value
+ .......................xyz Register file reads
+
+Tracking how much opcode overload,
+ 11111111111111110000000000000000
+ fedcba9876543210fedcba9876543210
+ ================================
+ bbb............................. Branching
+ ...rr........................... BRAM operation
+ .....sssstttuuuuvvvv............ Reg file ports
+ ....................??.......... Not enough space for DSP control
+ ......................iiiiiiiiii Trying for fixed 10-bit immediate
+ ================================
+ Notes,
+ Definitely need smarter opcode encoding
+
+Immediate,
+ 11111111111111110000000000000000
+ fedcba9876543210fedcba9876543210
+ ================================
+ ......................iiiiiiiiii Trying for fixed 10-bit immediate
+ ================================
+ Notes,
+ Wanted 10-bit to hit full BRAM absolute address
+
+DSP control,
+ 11111111111111110000000000000000
+ fedcba9876543210fedcba9876543210
+ ================================
+ Notes,
+ Lots of control bits to get correct in here
+
+DSP inputs,
+ 11111111111111110000000000000000
+ fedcba9876543210fedcba9876543210
+ ================================
+ ................................ Input a 24-bits
+ ................................ a=signExtend(i)
+ ................................ a=y[32:18]
+ ................................ a=f[32:18]
+ ................................ a=f
+ ................................ a=y
+ ................................ a=z
+ ================================
+ ................................ Input b 18-bits
+ ...............................0 b=i
+ ..............................?? b=y
+ ..............................?? b=f
+ ================================
+ ................................ Input c 32-bits
+ ...............................? c=f
+ ...............................? c=z
+ ================================
+ Notes,
+ Might be able to have a be modal based on DSP control (for mul vs a:b cases)
+ Or maybe merge for some cases?
+ The a=signExtend(i) case is needed for signed (a:b)=i
+ Should c=z be z or something else
+ Believe f probably should be an input into b for '(a:b) op c' case
+ Culled, a=i, as i is too small for top bits of a:b, and mul is associative
+ Culled, b=p, as it is better to just forward in these cases
+ Culled, c=p, hoping forwarding if control bits covers this
+
+Branch control,
+ 11111111111111110000000000000000
+ fedcba9876543210fedcba9876543210
+ ================================
+ .............................??? No branch
+ .............................??? Return
+ .............................??? Call to p
+ .............................??? Switch to/from MessageHandler/Program
+ .............................??? Conditional jump if p<0
+ .............................??? Conditional jump if p>=0
+ .............................??? Call
+ .............................??? Jump
+ ================================
+ Notes,
+ Likely not getting branch control under 3-bits
+ At least without multiple instruction forms
+ In theory this enables "free" branching
+ Won't have p==0, as don't want to turn on pattern detector
+ Work through computed branch targets cases again ...
+
+Register file ports,
+ 11111111111111110000000000000000
+ fedcba9876543210fedcba9876543210
+ ================================
+ .................ssss........... Store to any register
+ .....................ttt........ BRAM address registers (limited to first 8 for space)
+ ........................uuuuvvvv Load from any register
+ ================================
+ Notes,
+ This is the most trouble in opcode encoding size
+ Dropping to 16 entries from 32
+ Assuming want freer context switch to handle message
+ First 16 used during normal execution
+ Later 16 used during handle message
+
+BRAM operation bits,
+ 11111111111111110000000000000000
+ fedcba9876543210fedcba9876543210
+ ================================
+ ..............................00 f=[x]
+ ..............................01 f=[x^i]
+ ..............................10 [x]=p
+ ..............................11 [x^i]=p
+ ================================
+ Notes,
+ This is all that is needed to keep the one BRAM port busy
+ Direct mapping to register port t (1st read port) to save size
+ Only supporting storing from DSP output p
+ If one is going to store, store when generated
+ If need to store later just load back into the DSP (nop)
+ Need separate [x] case because imm may be used for something else
+ Culled, nop (just going to burn power for load regardless if needed)
+ Culled, f=[i], [i]=p, because [x^i] can load x=0
+
+
+===================
+ SHIFTING ISSUES
+===================
+Needs more thought ...
+
+Have the following pipeline options built in the DSP
+ (25-bit a * 18-bit b) + 48-bit c
+ ((25-bit a * 18-bit b) + 48-bit c) << 17
+
+Usage cases,
+ Address math,
+ This is 'base + index * stride', so use 'a*b+c'
+ Index limited to 16 M
+ Base not limited
+ Bitfield,
+ Included in via variable-bit BRAM access
+ Bit arrays,
+ Expanded to later section
+ Divides
+ Needs some thought ...
+
+Shifter will do at most 24-bit integers,
+ Top bits get sign-extended (likely not the opcode space for unsigned)
+ Using 24 to keep with quad LUT alignment
+
+Shifts for 16-bit,
+ __for_shift_>>__ __for_shift_<<__
+ 1111111111111111 0000000000000000
+ fedcba9876543210 fedcba9876543210 mul << >>
+ ================ ================ ======== == ==
+ ................ fedcba9876543210 00000001 0 10
+ ...............f edcba9876543210. 00000002 1 f
+ ..............fe dcba9876543210.. 00000004 2 e
+ .............fed cba9876543210... 00000008 3 d
+ ............fedc ba9876543210.... 00000010 4 c
+ ...........fedcb a9876543210..... 00000020 5 b
+ ..........fedcba 9876543210...... 00000040 6 a
+ .........fedcba9 876543210....... 00000080 7 9
+ ........fedcba98 76543210........ 00000100 8 8
+ .......fedcba987 6543210......... 00000200 9 7
+ ......fedcba9876 543210.......... 00000400 a 6
+ .....fedcba98765 43210........... 00000800 b 5
+ ....fedcba987654 3210............ 00001000 c 4
+ ...fedcba9876543 210............. 00002000 d 3
+ ..fedcba98765432 10.............. 00004000 e 2
+ .fedcba987654321 0............... 00008000 f 1
+ fedcba9876543210 ................ 00010000 10 0
+
+
+==============
+ BIT ARRAYS
+==============
+Maybe best to just use the 8-bit BRAM window
+ Keeps with-in the range of immediate for AND mask
+
+Emulation without special hardware,
+ Algorithms,
+ Extract lowest bit set .................... x & -x
+ Get mask up to lowest bit set ............. x ^ (x - 1)
+ Reset lowest bit set ...................... x & (x - 1)
+ =========================================== ============
+ Fill from lowest clear bit ................ x & (x + 1)
+ Isolate lowest clear bit and complement ... ~x & (x + 1)
+ Mask from lowest clear bit ................ x ^ (x + 1)
+ Mask from trailing zeros .................. ~x & (x - 1)
+ =========================================== ============
+ Isolate lowest clear bit .................. x | ~(x + 1)
+ Set lowest clear bit ...................... x | (x + 1)
+ Fill from lowest set bit .................. x | (x - 1)
+ Isolate lowest set bit and complement ..... ~x | (x - 1)
+ Inverse mask from trailing ones ........... ~x | (x + 1)
+
+Common algorithms,
+ Bit insert
+ Bit extract
+ Popuplation count
+ Output is 3 bits
+ Requires 8:1 function, or 2 LUTs/bit, 6 LUTs in hardware
+ Count leading zeros
+ 6 LUTs in hardware
+ Count trailing zeros
+ 6 LUTs in hardware
+
+Todo,
+ Look through De Bruijn Sequence based stuff again
+
+
+
+================
+ FORTH HYBRID
+================
+Dual stack machine with register file
+Core functional units,
+
+ 16-bit x 8-entry return stack (1 port)
+ 32-bit x 8-entry data stack (1 port)
+ 32-bit x 8-entry register file (2 ports, port 0 for DSP input, port 1 for BRAM address)
+ 32-bit x 1024-entry BRAM (2 ports, port 0 for instruction fetch, port 1 for data)
+
+
+===================
+ EXECUTION MODEL
+===================
+4 threads/core of execution running with guaranteed round-robin scheduling
+Instructions are VLIW style with a fixed logical ordered set of operations
+
+ Order Operation
+ ===== =========
+ 1st mux inputs for DSP ............ uses loads from prior instruction
+ 2nd DSP execution .................
+ 3nd load/store to REG and BRAM .... can store DSP result
+ 4th branch ........................ can branch to DSP result
+
+
+=====================
+ PHYSICAL PIPELINE
+=====================
+Designing under the following constraints,
+
+ Loads from RAMs are not used until next stage
+ Each stage does only one LUT or ADD both of which can be vertically chained
+ DSP is fully pipelined
+
+Outline,
+
+ DSP DSP DSP DSP DSP DAT ADR BLK BLK
+ Stage A B MUL C ADD P REG REG OUT RAM PC INS
+ ===== === === === === === === === === === === ===
+ 0 lut lut @
+ 1 mul reg lut
+ 2 add lut add
+ 3 lut @! lut @! lut @
+ ----- --- --- --- --- --- --- --- --- --- --- ---
+
+ DSP A B .... DSP input a and b arguments
+ DSP MUL .... DSP mul stage
+ DSP C ...... DSP input c argument
+ DSP ADD .... DSP add/op stage
+ DAT REG .... Register file data load/store
+ ADR REG .... Register file address register to BRAM address translation
+ BLK OUT .... BRAM data construct write value
+ BLK RAM .... BRAM data load/store
+ PC ......... Update program counter
+ INS ........ Fetch next instruction
+
+
+============================
+ CURRENT LUT BUDGET USAGE
+============================
+Budget is 400 LUTs/core, adding as design is roughed out,
+
+ LUTs % usage
+ ==== === =====
+ 32 8 register file (4x 8-LUT SLICEM 32-entry x 8-bit 2 port RAM)
+ 16 4 data stack (2x 8-LUT SLICEM 32-entry x 16-bit 1 port RAM)
+ 8 2 return stack ( 8-LUT SLICEM 32-entry x 16-bit 1 port RAM)
+ ---- --- -----
+ 54 DSP b input
+ DSP a input
+ DSP c input
+ 18 5 BRAM address generation
+ 34 BRAM output generation
+ 38 10 program counter
+ ==== === =====
+ 200 total
+
+
+========
+ TODO
+========
+Make sure to register all inputs required to generate a b and c
+
+
+===========================
+ DSP B INPUT
+===========================
+Need to also unpack BRAM load options so this gets expensive
+Placement in pipeline,
+
+ stage action
+ ===== ======
+ 0 LUT DSP b input
+ 1
+ 2 register pre-translated address for stage 0 of next cycle
+ 3 register pre-translated address for stage 0 of next cycle
+
+The b input expanded with unpack options, and control bits,
+
+ fedcba9876543210 n LUT input count
+ ================ = ===============
+ <-----iiiiiiiiii i 1-bit
+ tttttttttttttttt t 1-bit
+ dddddddddddddddd d 1-bit
+ ================
+ aaaaaaaaaaaaaaaa f 4-bits for MSB 8-bits of output
+ bbbbbbbbbbbbbbbb 7-bits for 2nd LSB 4-bits of output
+ cccccccccccccccc 15-bits for LSB 4-bits of output
+ 00000000dddddddd
+ 00000000eeeeeeee
+ 00000000ffffffff
+ 00000000gggggggg
+ 000000000000hhhh
+ 000000000000iiii
+ 000000000000jjjj
+ 000000000000kkkk
+ 000000000000llll
+ 000000000000mmmm
+ 000000000000nnnn
+ 000000000000oooo
+ ================
+ xxxxxxxxxxxxxxxx needs 2-bit opcode control
+ xxxxxxxxxxxxxxxx needs 2-bit MSB of pre-translate address
+ xxxxxxxx........ needs 1-bit LSB of pre-translate address
+ ........xxxx.... needs 2-bit LSB of pre-translate address
+ ............xxxx needs 3-bit LSB of pre-translate address
+ ================
+ xxxxxxxx........ 12:1 function (2 LUT/bit) x 8-bit = 16 LUT
+ ........xxxx.... 16:1 function (4 LUT/bit) x 4-bit = 16 LUT
+ ............xxxx 25:1 function (4 LUT/bit) x 4-bit = 16 LUT
+
+LUT area estimate,
+
+ LUTs usage
+ ==== =====
+ 48 generate b
+ 6 to register 5-bits x 2 stages of pre-translate address (rounded up)
+ ---- -----
+ 54 total
+
+
+===========================
+ DSP A INPUT
+===========================
+Placement in pipeline,
+
+ stage action
+ ===== ======
+ 0 LUT DSP a input
+ 1
+ 2
+ 3
+
+LUT area estimate,
+
+ LUTs usage
+ ==== =====
+ ---- -----
+ total
+
+
+===========================
+ DSP C INPUT
+===========================
+Placement in pipeline,
+
+ stage action
+ ===== ======
+ 0 LUT DSP c input
+ 1 register c
+ 2
+ 3
+
+LUT area estimate,
+
+ LUTs usage
+ ==== =====
+ ---- -----
+ total
+
+
+===========================
+ BRAM ADDRESS GENERATION
+===========================
+Supports the feature of variable-bit width windows into the 4KB of ram
+Placement in pipeline,
+
+ stage action
+ ===== ======
+ 0 fetch base address from register file
+ 1 optionally XOR immediate
+ 2 translate into BRAM address
+ 3
+
+Implementation requires XOR control to be single bit in opcode
+
+BRAMs always in 32-bit port mode,
+
+ fedcba9876543210
+ ================
+ .xxxxxxxxxx00000 - requires 10-bit address
+
+Address register,
+
+ fedcba9876543210 access
+ ================ ======
+ 00....xxxxxxxxxx 1024 x 32-bit
+ 01...xxxxxxxxxxx 2048 x 16-bit
+ 10..xxxxxxxxxxxx 4096 x 8-bit
+ 11.xxxxxxxxxxxxx 8192 x 4-bit (supported for read only)
+
+Address register value to BRAM address translation
+ Uses a 6:1 function for each bit,
+
+ bits meaning
+ ==== =======
+ 4 address shifted left {0,1,2,3} bits
+ 2 the 'fe' address bits
+
+LUT area estimate,
+
+ LUTs usage
+ ==== =====
+ 8 optional XOR (16-bits at 2-bits per LUT), rounding up for ending register
+ 10 translate (10-bits x 1 LUT)
+ ---- -----
+ 18 total
+
+
+===========================
+ BRAM OUTPUT GENERATION
+===========================
+Shifts DSP p output for store, and compute byte write mask
+Placement in pipeline,
+
+ stage action
+ ===== ======
+ 0
+ 1
+ 2
+ 3 LUT new output here
+
+Permutations (showing address and byte write mask for store),
+
+ 11111111111111110000000000000000
+ fedcba9876543210fedcba9876543210
+ ================================
+ aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa - 32-bit adr=00....xxxxxxxxxx write=1111
+ ................bbbbbbbbbbbbbbbb - 16-bit adr=01...xxxxxxxxxx0 write=0011
+ cccccccccccccccc................ - 16-bit adr=01...xxxxxxxxxx1 write=1100
+ ........................dddddddd - 8-bit adr=10..xxxxxxxxxx00 write=0001
+ ................eeeeeeee........ - 8-bit adr=10..xxxxxxxxxx01 write=0010
+ ........ffffffff................ - 8-bit adr=10..xxxxxxxxxx10 write=0100
+ gggggggg........................ - 8-bit adr=10..xxxxxxxxxx11 write=1000
+ ............................hhhh - 4-bit adr=11.xxxxxxxxxx000
+ ........................iiii.... - 4-bit adr=11.xxxxxxxxxx001
+ ....................jjjj........ - 4-bit adr=11.xxxxxxxxxx010
+ ................kkkk............ - 4-bit adr=11.xxxxxxxxxx011
+ ............llll................ - 4-bit adr=11.xxxxxxxxxx100
+ ........mmmm.................... - 4-bit adr=11.xxxxxxxxxx101
+ ....nnnn........................ - 4-bit adr=11.xxxxxxxxxx110
+ oooo............................ - 4-bit adr=11.xxxxxxxxxx111
+
+Shift value for store,
+ Requires 3:1 MUX per bit, 32 LUTs
+
+Generate write enable for store,
+ Requires same 4-bits per function,
+ 2 lower address bits
+ 2 upper address bits
+ 2 LUTs (5:1 function sharing inputs, 2 outputs per LUT)
+
+LUT area estimate,
+
+ LUTs usage
+ ==== =====
+ 32 shift value for store
+ 2 generate write enable
+ ---- -----
+ 34 total
+
+
+===================
+ PROGRAM COUNTER
+===================
+10-bit program counter (PC)
+Only lower 8-bits of PC increment on linear execution
+ Requires only an 8-bit PC+1 computation (one slice)
+
+Placement in pipeline,
+
+ stage action
+ ===== ======
+ 0 register
+ 1 register
+ 2 increment PC
+ 3 LUT new PC based on DSP p output and instruction opcode
+
+PC function inputs per output bit (map to 13:1 function at 2 LUTs/bit),
+
+ bits meaning
+ ==== =======
+ 1 next PC if not branching (computed in prior stage)
+ 1 top of return stack
+ 1 immediate absolute branch address
+ 1 DSP P output register (computed branch target in prior clock)
+ 1 DSP P output register sign bit (for conditional branch)
+ 3 bits from instruction opcode
+
+LUT area estimate,
+
+ LUTs usage
+ ==== =====
+ 20 13:1 function for next 10-bit PC computation including instruction decode
+ 8 PC+1 adder for 8 lower bits of PC
+ 10 for 2 stage registers (2-bits/LUT)
+ ---- -----
+ 38 total
+
+
+========================
+ INSTRUCTION PIPELINE
+========================
+Todo, remember to count cost to pipeline opcode bits through stages
+
+
+
+=========
+
+ NOTES
+
+=========
+
+=================================================
+ ADDRESS REGISTER XOR INSTEAD OF ADD IMMEDIATE
+=================================================
+Planning on [address ^ immediate] addressing
+ This removes an adder from the design
+
+XOR is the same as [address + immediate] for an n-bit immediate
+ When lower n-bits of address are zero
+ Means data must be aligned to the nearest pow2 of maximum immediate offset
+
+Using the following terms,
+ ggggggoooo
+ g = group bits (address bits choose the group of data)
+ o = offset bits (address bits are zero, immediate chooses element in group)
+
+For bits in address which are not cleared (ie the group bits),
+ Setting bits in the immediate results in accessing a neighbor group
+ Regardless of the starting group in the address register
+ It is possible to roll through all aligned groups
+ But ordering is different based on starting group address
+ Example of group bits for address crossed with immediate
+
+ 00 01 10 11
+ +-------------
+ 00 | 00 01 10 11
+ 01 | 01 00 11 10
+ 10 | 10 11 00 01
+ 11 | 11 10 01 00
+
+
+===================================
+ BRAM VARIABLE BIT-WIDTH WINDOWS
+===================================
+Trying to support transparent pack/unpack of variable bit-widths from BRAM
+ Want zero impact to ISA, no special instructions
+ Instead dividing address range into windows of different bit-widths
+ Each address range addresses at a multiple of the bit-width
+ Effectively the high bits of address choose the bit-width
+
+Store path limited to {8,16,32}-bit
+ Only using BRAM byte write mask to avoid any {read, modify, write}
+
+Fixed signed vs unsigned configuration,
+
+ size choice
+ ====== ======
+ 32-bit signed (but doesn't matter)
+ 16-bit going to go with signed (needed for vector or audio)
+ 8-bit unsigned (keeps implementation simple)
+ 4-bit unsigned for sure (sprites?)
+
+
+==========================================
+ WORKING THROUGH OPTIONS DSP OPERATIONS
+==========================================
+Opcode forms,
+
+ p = c op ((a << 16) + unsigned(b))
+ p = c + (a * b)
+ p = c - (a * b)
+
+Where op can be the following,
+
+ and .....
+ nand ....
+ nor .....
+ not .....
+ or ......
+ xnor ....
+ xor .....
+
+Where the following can also be applied,
+
+ extra set c bit -1 to 1 (for rounding)
+ (a * b) can be forced to zero (nop)
+ ((a << 16) + unsigned(b)) can be forced to zero (nop)
+ ((a << 16) + unsigned(b)) can be forced to all ones
+
+
+===============================================
+ WORKING THROUGH OPTIONS FOR A,B,C DSP INPUT
+===============================================
+DSP inputs (as they appear in the core),
+
+ 24-bit a
+ 16-bit b
+ 40-bit c
+
+Possible inputs,
+
+ 10-bit immediate
+ 16-bit top of return stack
+ 32-bit top of data stack
+ 32-bit register file load (from prior instruction)
+ 32-bit BRAM load (from prior instruction)
+ 40-bit DSP p output
+
+
+====================
+ FAST ABS MIN MAX
+====================
+Simple design exercise to think through DSP issues
+
+These need to work on the 40-bit accumulator without precision loss
+ So using multiply stage is out
+
+Min and max, where a is the accumulator, and b is the limit,
+
+ min(a, b) = ((a - b) & ((a - b) < 0 ? ~0 : 0)) + b
+ max(a, b) = ((a - b) & ((a - b) < 0 ? 0 : ~0)) + b
+
+Want to be able do the following,
+
+ acc -= b;
+ acc = acc < 0 ? acc : 0; // want to fold this into prior operation
+ acc += b;
+
+Have to either LUT or register p in stage 3,
+ Could LUT p to zero if signed or unsigned based on control bit
+ This works out to 2-bits/LUT (pair of 5:1 functions with same input)
+ So 20 LUTs total (same as just registering)
+ Plus likely need to decode control and enable from opcode in prior pass
+
+ inputs
+ ------
+ 2 p bits
+ 1 p sign bit (might want the overflow sign bit?)
+ 1 enable bit
+ 1 signed or unsigned control bit
+
+This enables min and max to work in 2 instructions without branching
+
+Absolute value,
+
+ abs(a) = max(a, -a)
+
+Does this make the case for,
+
+ expanding data stack to 40-bit (to match accumulator)?
+ reducing accumulator to 36-bit, or even 32-bit?
+
+Operation,
+
+ push copy of acc; // want to fold into start of next op
+ acc += acc; acc = acc < 0 ? 0 : acc;
+ acc -= pop;
+
+Using top of data stack for DSP input means it must be pre-registered
+ That register could be 40-bit until it gets actually stored on stack
+ Want a bit which marks if should consume data stack
+
+Ideally push to happen before the first add (included in that opcode)
+ Stage 0 : must save top to stack RAM
+ Stage 1 : set top to p
+ Todo, think through when c is computed again
+
+Data stack top is going to be expensive
+ 40-bit : minimum 80 LUTs
+ 32-bit : minimum 64 LUTs
+
+Time to rethink ...
+
+
+Related
+