multiformats / multiformats/py-multihash

No `Cast()` or `MHFromBytes()` utility functions

Open
#56 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
17
Forks
17
Avg merge
1h 5m
Merged PRs (30d)
3

Description

Go provides Cast(buf []byte) (Multihash, error) for validating and casting raw bytes to a Multihash, and MHFromBytes(buf []byte) (int, Multihash, error) for reading a multihash from the beginning of a buffer with bytes-consumed tracking. Python has decode() but no direct equivalents.

Problem

Go's multihash.go:

// Cast casts a buffer onto a multihash, and returns an error if it does not work.
func Cast(buf []byte) (Multihash, error) {
    _, err := Decode(buf)
    if err != nil { return nil, err }
    return Multihash(buf), nil
}

// MHFromBytes reads a multihash from the given byte buffer,
// returning the number of bytes read and the multihash.
func MHFromBytes(buf []byte) (int, Multihash, error) {
    nr, _, _, err := readMultihashFromBuf(buf)
    if err != nil { return 0, nil, err }
    return nr, Multihash(buf[:nr]), nil
}

MHFromBytes is particularly useful when parsing a stream that contains a multihash followed by other data — it tells you exactly how many bytes the multihash consumed.

Python's decode() always consumes the entire input and doesn't return bytes consumed.

Proposed Solution

Add both functions:

def cast(buf: bytes) -> bytes:
    """Validate and return multihash bytes. Raises ValueError if invalid."""
    decode(buf)  # Validates
    return buf

def from_bytes_with_length(buf: bytes) -> tuple[int, "Multihash"]:
    """Read a multihash from the beginning of buf.

    Returns (bytes_consumed, Multihash).
    Unlike decode(), this handles buffers with trailing data.
    """
    buffer = BytesIO(buf)
    code = varint.decode_stream(buffer)
    if not is_valid_code(code):
        raise ValueError(f"Unsupported hash code {code}")
    length = varint.decode_stream(buffer)
    digest = buffer.read(length)
    if len(digest) != length:
        raise ValueError(f"Insufficient data: expected {length} bytes, got {len(digest)}")
    bytes_consumed = buffer.tell()
    name = constants.CODE_HASHES.get(code, code)
    return bytes_consumed, Multihash(code=code, name=name, length=length, digest=digest)

Export from __init__.py and add tests.

Related

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reading the existing decode() implementation and init.py exports, then inspect the related parsing and varint handling used by the library. Add the two utilities with validation and consumed-byte behavior for buffers with trailing data, export them, and add tests covering valid, invalid, truncated, and trailing-data inputs.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
backend
Issue type
Feature
Difficulty
3/5
Estimated time
1-2 days
Activity status
Quiet
Clarity
Clearly specified
Newbie friendliness
68/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.