[Feature request] Use the bytes of an opcode in numeric expressions
- Dominant language
- C++
- Stars
- 1.6k
- Forks
- 193
- Avg merge
- 22h 17m
- Merged PRs (30d)
- 26
Description
`LOAD` blocks exist for whole chunks of RAM code. However, self-modifying code usually needs more fine-grained control, writing and rewriting individual bytes, not just copying N bytes from ROM to RAM. Here's an example:
```
; Build a function to write pixels in hAppendVWFText.
; - nothing: or [hl] / ld [hld], a / ld [hl], a / ret
; - invert: xor [hl] / ld [hld], a / ld [hl], a / ret
; - opaque: or [hl] / ld [hld], a / ret
; - invert+opaque: xor [hl] / ld [hld], a / ret
ld hl, hAppendVWFText
bit VWF_INVERT_F, b
ld a, $ae ; xor [hl]
jr nz, .invert
ld a, $b6 ; or [hl]
.invert
ld [hli], a
ld a, $32 ; ld [hld], a
ld [hli], a
bit VWF_OPAQUE_F, b
jr nz, .opaque
ld a, $77 ; ld [hl], a
ld [hli], a
.opaque
ld [hl], $c9 ; ret
```
And another:
```
LD_A_FFXX_OP EQU $f0
JR_C_OP EQU $38
JP_C_OP EQU $da
LD_B_XX_OP EQU $06
RET_OP EQU $c9
RET_C_OP EQU $d8
DEC_C_OP EQU $0d
JR_NZ_OP EQU $20
LD_A_HLI_OP EQU $2a
LD_C_XX_OP EQU $0e
ADD_A_OP EQU $87
CopyBitreeCode:
ld a, DEC_C_OP
ld [hli], a
ld a, JR_NZ_OP
ld [hli], a
ld a, 3
ld [hli], a
ld a, LD_A_HLI_OP
ld [hli], a
ld a, LD_C_XX_OP
ld [hli], a
ld a, 8
ld [hli], a
ld a, ADD_A_OP
ld [hli], a
ret
```
`rgbasm` can already encode instructions as bytes, so it would be convenient to have a syntax allowing those bytes as part of usual numeric expressions. A few ideas:
- `<[ nop ]>` (inspired by HTML CDATA)
- `'nop'` (not currently in use)
- `OPCODE(nop)` (like `HIGH(bc)` or `DEF(Symbol)`)
(I think `OPCODE` would be clearest and fit in best with existing `rgbasm` syntax; too much meaningful punctuation ends up like Perl.)
Some tricky details:
- How to handle multi-byte operations. `rl h` acts like `db $CB, $14`; this could be pretty easily handled with `HIGH` and `LOW`, like `ld a, LOW(OPCODE(rl h)) / ld [hli], a / ld [hl], HIGH(OPCODE(rl h))`. (Little-endian order to be consistent, so `dw OPCODE(rl h)` would act like plain `rl h`.) Three-byte ones would still be feasible, if less convenient: do `def op = OPCODE(ld hl, $abcd)`, then work with `LOW(op)` ($21), `HIGH(op)` ($cd), and `op>>16` ($ab).
- How to handle relative jumps. Evaluating the relative position of a label would be confusing and, I expect, not useful. People are more likely to want absolute jump distances, which currently get expressed with `@`: e.g. jumping ahead 5 bytes without a label is done as `jr (@ + 2) + 5`. Here, `OPCODE(jr 5)` could evaluate to `$0518`, or `OPCODE(jr z, $ff)` to `$ff28` (aka the notional "`rst z, $38`"). Or just disallow the destination here, so `OPCODE(jr) == $18` and `OPCODE(jr z) == $28`.
However this is done, it should reuse the parser's regular opcode handling, without needing two separate paths. I don't expect that to be a problem.
Contributor guide
Research direction
Start by tracing rgbasm's regular opcode parser and the numeric-expression parser; no specific files or tests are named in the issue. Compare the proposed OPCODE forms with existing HIGH, LOW, and DEF handling, then establish behavior for multi-byte instructions and relative jumps before implementing tests for the accepted syntax.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- compilers
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 30/100