Repository navigation
py-multibase ↔ go-multibase Feature Parity Analysis #1362
sumanjeet0012
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
py-multibase Feature Parity Analysis
Table of Contents
Executive Summary
The Python implementation (
py-multibase) supports a solid set of 24 encodings including some that Go does not implement (base8, base10, base32z). The Go implementation (go-multibase) is leaner with 21 encodings but has stronger tooling: a typedEncoderwithEncoderByNamelookup (by name or single-character prefix), a CLI transcoding tool, official spec-compliance tests, benchmarks, and fuzz testing.Encoder/EncoderByNameEncoderlacks name-or-prefix lookupMustNewEncoderreturn_encoding=Truemultibase-conv)multibase.csvor test vectorsArchitecture Comparison
Go Architecture
Dependencies:
mr-tron/base58,multiformats/go-base32,multiformats/go-base36— all optimised native Go libraries.Python Architecture
Dependencies:
python-baseconv(generic arbitrary-base conversion),morphys(bytes/string helpers),six(Python 2/3 compat — no longer needed for Python 3.10+).Feature-by-Feature Gap Analysis
3.1 Supported Encodings
EncodingToStr)ENCODINGS)'0''7''9''f''F''b''B''c''C''v''V''t''T''h''k''K''z''Z''m''M''u''U''🚀''R'Summary: Python supports 24 encodings; Go supports 21. Python leads on base8, base10, base32z. Go defines Base45 as a constant but doesn't implement it; Python doesn't mention it at all.
3.2 Core Encode / Decode API
Encode(base, data)→ stringEncode(Encoding, []byte) (string, error)encode(encoding, data) → bytesDecode(data)→ encoding + bytesDecode(string) (Encoding, []byte, error)— always returns encodingdecode(data) → bytes(encoding only withreturn_encoding=True)IndexErrororInvalidMultibaseStringErrorstringbytesEncoding(int constant)str(name)3.2.1 Empty-string / zero-length handling
Go explicitly checks:
Python has no equivalent guard — an empty
b""will reachget_codecwhich triesdata[:1]and either raisesIndexError(unlikely since slicingb""[:1]yieldsb"") or falls through to aKeyErrorwrapped inInvalidMultibaseStringError. This works but is less explicit.3.2.2 Decode always returns encoding
Go's
Decodesignature always returns the detected encoding as the first return value. Python makes this opt-in:Gap: Python's default hides the encoding. For Go parity, the default should arguably return both.
3.3 Encoder Type
EncoderEncoderNewEncoder(base Encoding)EncoderByName(str)EncoderByName("f")→ base16 encoderMustNewEncoder(base)Encode(data)methodstring(no error, pre-validated)bytesEncoding()accessorEncoder.Encoding() Encodingencoder.encoding(string name, not code)Key missing feature:
EncoderByNamewhich accepts either the human-readable name ("base16") or the single-character multibase prefix ("f") and returns an Encoder. This is very useful for CLI tools and configuration-driven encoding.3.4 Encoding Lookup & Introspection
EncodingToStrmap (code → name)ENCODINGS_LOOKUP(code → Encoding)Encodingsmap (name → code)ENCODINGS_LOOKUP(name → Encoding)is_encoded(data)is_encoding_supported(name)list_encodings()get_encoding_info(name)get_codec(data)→ codec infoDecoder/ComposedDecoderdecode(data, return_encoding=True)Python has significantly more introspection and utility APIs. Go is more minimal and expects users to work directly with the
EncodingToStr/Encodingsmaps.3.5 Error Handling
ErrUnsupportedEncodingUnsupportedEncodingErrorfmt.Errorf("cannot decode ... zero length string")fmt.Errorf("unsupported multibase encoding: %d")UnsupportedEncodingErrorfmt.Errorf("empty multibase encoding")base256emojiCorruptInputError{index, char}ValueError/DecodingErrorMultibaseError(base)InvalidMultibaseStringErrorDecodingError3.6 CLI Tool
multibase-convmultibase-conv <new-base> <multibase-str>...The Go CLI is simple but useful for quick transcoding:
multibase-conv z f796563206d616e692021 # base16 → base58btc3.7 Testing, Spec Compliance & Benchmarks
encodedSamplesmap (all encodings, sample text)TEST_FIXTURES(many encodings, varied inputs)multibase.csvTestSpec— validates encoding names/codes against official tablespec/tests/*.csvTestSpecVectors— encode + decode against official vectorsFuzzDecodeseeded with spec vectorsTestBase256EmojiAlphabetTestBase256EmojiUniqTestEncoder,TestInvalidCode,TestInvalidNametest_encoder_classTestMapFeature Matrix Summary
EncoderByName(name_or_prefix)Encoderconstructible by code constantEncoder.encodingreturns code (not just name)MustNewEncoder(strict factory)decode()returns encoding by defaultdecode()base256emojiCorruptInputErrorwith indexmultibase-conv)multibase.csv)spec/tests/*.csv)FuzzDecode)sixdependencyversion.jsonPhased Roadmap
Phase 1 – Correctness Fixes & Core API Gaps
Goal: Fix correctness issues and add the most important missing APIs.
Estimated effort: 1 week
1.1 Add
EncoderByNamefunctionPort Go's
EncoderByNamewhich accepts either a human-readable name ("base16") or a single-character multibase prefix ("f"):1.2 Make
Encoderconstructible by code constant1.3 Add
encoding_codeproperty toEncoder1.4 Explicit empty-string check in
decode()1.5 Add
base256emojicorrupt-input error with indexPhase 2 – Encoder Parity & API Enrichment
Goal: Full parity with Go's
Encodertype and add useful utilities.Estimated effort: 3–5 days
2.1 Add
transcode()utilityMirroring Go's
multibase-convfunctionality:2.2 Add
MustNewEncoderequivalent (strict factory)2.3 Add case-insensitive base32 decode verification
Verify that the Python
Base32StringConverterproperly handles case-insensitive input. Go usesNewEncodingCIwhich accepts both upper and lower. The Python implementation usesself.digits.index(byte_)which is case-sensitive — this meansdecode("bIRSWGZLOORZGC3DJPJSSAZLWMVZHS5DINFXGOIJBEE")would fail forbase32(lowercase digits).Fix: Make base32 decode case-insensitive regardless of the encoding variant:
2.4 Add
base64url/base64urlpadround-trip verificationVerify that URL-safe base64 variants produce correct output matching Go's
base64.RawURLEncoding/base64.URLEncoding.2.5 Remove stale
sixdependencyThe
pyproject.tomllistssix>=1.10.0,<2.0as a dependency. Since the project requires Python ≥3.10,sixis unnecessary. Remove it and verifymorphysdoesn't re-introduce it transitively.Phase 3 – Spec Compliance & Testing Hardening
Goal: Add official spec compliance tests and harden the test suite to match Go's rigour.
Estimated effort: 1–2 weeks
3.1 Add official multibase spec as a git submodule
This provides:
spec/multibase.csv— the official encoding tablespec/tests/*.csv— official test vectors3.2 Implement
TestSpec— validate encoding table3.3 Implement
TestSpecVectors— official test vectors3.4 Add round-trip tests with random data
Port Go's comprehensive round-trip test:
3.5 Add invalid-input rejection tests
3.6 Add case-insensitive decode tests
3.7 Add base256emoji alphabet tests
3.8 Add encoding map round-trip test
3.9 Add fuzz testing
Using
pytest-fuzzor hypothesis:Phase 4 – CLI Tool & Performance
Goal: Provide a CLI tool matching Go's
multibase-convand add benchmarks.Estimated effort: 1 week
4.1 Create CLI entry point
Register as console script in
pyproject.toml:4.2 Add benchmark suite
Using
pytest-benchmark:4.3 Add
version.json{ "version": "v2.0.0" }4.4 Performance audit: Replace
python-baseconvfor hot pathsThe current
BaseStringConverterusespython-baseconvwhich does arbitrary-base conversion via integer math. For base2 and other power-of-2 bases, direct bit manipulation (like Go'sbase2.go) would be significantly faster. Consider:bin()/int(s, 2)directly (matching Go'sstrconv.ParseInt(s, 2, 0))binascii.hexlify/bytes.fromhex(already done for encode, check decode)base64.b32encode/base64.b64encodewith appropriate alphabetsbase36or custom implementationbase58PyPI package (Bitcoin-optimised)Appendix – Encoding Implementation Details
A.1 Base32 Case Sensitivity
'b')NewEncodingCI)digits.index())'B')NewEncodingCI)'c')'v')A.2 Base64 Implementation Differences
base64.RawStdEncoding(stdlib)Base64StringConverterwith bit manipulationbase64.StdEncoding(stdlib)Base64StringConverterwith padding logicbase64.RawURLEncoding(stdlib)Base64StringConverterbase64.URLEncoding(stdlib)Base64StringConverterGo uses the standard library's battle-tested base64 implementations. Python uses a custom bit-manipulation approach. Both should produce identical output, but the Python implementation should be verified against the Go output for all edge cases.
A.3 Base58 Libraries
mr-tron/base58python-baseconv(genericBaseConverter)b58.BTCAlphabetENCODINGSb58.FlickrAlphabetENCODINGSA.4 Dependency Comparison
mr-tron/base58python-baseconvmultiformats/go-base32Base32StringConverter)multiformats/go-base36python-baseconvencoding/base64Base64StringConverter)morphyssix(unnecessary for 3.10+)A.5 Base8 and Base10 — Python-Exclusive
Python implements base8 and base10 using the generic
BaseStringConverter. Go defines the constants ('7'for base8,'9'for base10) but does not implement them — they are absent fromEncodingToStrand theEncode/Decodeswitch statements. This means Python is the only reference implementation for these encodings.Recommendation: Since these are part of the multibase spec but rarely used, document their Python-exclusive status clearly. Consider contributing test vectors to the official spec repository.
Summary Timeline
EncoderByName, code-constant constructor, empty-string guard, emoji error with indextranscode(), case-insensitive base32, removesix, URL-safe base64 verificationmultibase-convCLI, benchmark suite, performance audit for base2/16/32/64Total estimated effort: 4–5 weeks
Each phase is independently shippable. Phase 1 and 3 are the highest priority — Phase 1 fixes real API gaps, and Phase 3 ensures correctness against the official spec. Phase 2 and 4 are quality-of-life improvements.
All reactions