gbdev / gbdev/rgbds

[Feature request] Use the bytes of an opcode in numeric expressions

Open
#823 13 comments 0 reactions 0 assignees View on GitHub
enhancement rgbasm
Dominant language
C++
Stars
1.6k
Forks
193
Avg merge
22h 17m
Merged PRs (30d)
26

Description

`LOAD` blocks exist for whole chunks of RAM code. However, self-modifying code usually needs more fine-grained control, writing and rewriting individual bytes, not just copying N bytes from ROM to RAM. Here's an example:

```
; Build a function to write pixels in hAppendVWFText.
; - nothing: or [hl] / ld [hld], a / ld [hl], a / ret
; - invert: xor [hl] / ld [hld], a / ld [hl], a / ret
; - opaque: or [hl] / ld [hld], a / ret
; - invert+opaque: xor [hl] / ld [hld], a / ret
ld hl, hAppendVWFText
bit VWF_INVERT_F, b
ld a, $ae ; xor [hl]
jr nz, .invert
ld a, $b6 ; or [hl]
.invert
ld [hli], a
ld a, $32 ; ld [hld], a
ld [hli], a
bit VWF_OPAQUE_F, b
jr nz, .opaque
ld a, $77 ; ld [hl], a
ld [hli], a
.opaque
ld [hl], $c9 ; ret
```

And another:

```
LD_A_FFXX_OP EQU $f0
JR_C_OP EQU $38
JP_C_OP EQU $da
LD_B_XX_OP EQU $06
RET_OP EQU $c9
RET_C_OP EQU $d8

DEC_C_OP EQU $0d
JR_NZ_OP EQU $20
LD_A_HLI_OP EQU $2a
LD_C_XX_OP EQU $0e
ADD_A_OP EQU $87

CopyBitreeCode:
ld a, DEC_C_OP
ld [hli], a
ld a, JR_NZ_OP
ld [hli], a
ld a, 3
ld [hli], a
ld a, LD_A_HLI_OP
ld [hli], a
ld a, LD_C_XX_OP
ld [hli], a
ld a, 8
ld [hli], a
ld a, ADD_A_OP
ld [hli], a
ret
```

`rgbasm` can already encode instructions as bytes, so it would be convenient to have a syntax allowing those bytes as part of usual numeric expressions. A few ideas:

- `<[ nop ]>` (inspired by HTML CDATA)
- `'nop'` (not currently in use)
- `OPCODE(nop)` (like `HIGH(bc)` or `DEF(Symbol)`)

(I think `OPCODE` would be clearest and fit in best with existing `rgbasm` syntax; too much meaningful punctuation ends up like Perl.)

Some tricky details:

- How to handle multi-byte operations. `rl h` acts like `db $CB, $14`; this could be pretty easily handled with `HIGH` and `LOW`, like `ld a, LOW(OPCODE(rl h)) / ld [hli], a / ld [hl], HIGH(OPCODE(rl h))`. (Little-endian order to be consistent, so `dw OPCODE(rl h)` would act like plain `rl h`.) Three-byte ones would still be feasible, if less convenient: do `def op = OPCODE(ld hl, $abcd)`, then work with `LOW(op)` ($21), `HIGH(op)` ($cd), and `op>>16` ($ab).
- How to handle relative jumps. Evaluating the relative position of a label would be confusing and, I expect, not useful. People are more likely to want absolute jump distances, which currently get expressed with `@`: e.g. jumping ahead 5 bytes without a label is done as `jr (@ + 2) + 5`. Here, `OPCODE(jr 5)` could evaluate to `$0518`, or `OPCODE(jr z, $ff)` to `$ff28` (aka the notional "`rst z, $38`"). Or just disallow the destination here, so `OPCODE(jr) == $18` and `OPCODE(jr z) == $28`.

However this is done, it should reuse the parser's regular opcode handling, without needing two separate paths. I don't expect that to be a problem.

Contributor guide

Open the contributing guide

Research direction

Start by tracing rgbasm's regular opcode parser and the numeric-expression parser; no specific files or tests are named in the issue. Compare the proposed OPCODE forms with existing HIGH, LOW, and DEF handling, then establish behavior for multi-byte instructions and relative jumps before implementing tests for the accepted syntax.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
compilers
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.