Skip to content

Performance Evolution: The AST Revolution

Lodum's performance is the result of systematic experimentation and a fundamental shift in how Python serialization is handled. This document tracks our journey from a dynamic baseline to a high-performance, bytecode-compiled engine.

Performance Trajectory

Standardized on Linux (ubuntu-latest) hardware. Measurements in microseconds (us) per operation.

Scenario v0.1.0 (Initial) v0.2.0 (AST Core) v0.3.0 (Modular) Total Improvement
Simple Serialization 5.56 4.86 4.90 12% faster
Simple Deserialization 18.82 12.05 12.47 34% faster
Complex Serialization 10.95 7.79 7.90 28% faster
Complex Deserialization 37.31 25.99 26.35 29% faster
Nested Serialization 25.58 21.54 21.69 15% faster
Nested Deserialization 117.81 72.63 74.44 37% faster

Core Optimization Experiments

Our optimization strategy focused on eliminating the "Interpreter Tax"—the overhead of generic loops and repeated runtime inspections.

1. Pre-Resolved Handlers (Experiment 2)

Originally, every field access triggered a cache lookup and type introspection. - The Optimization: We moved handler resolution to compile-time. The generated function now has direct references to the handlers it needs. - Impact: Achieved a 65% reduction in dumping overhead.

2. Fast Raw-Dict Access (Experiment 5)

Traditional loaders often wrap input data in several layers of abstraction. - The Optimization: We implemented a "Fast Path" that allows the generated bytecode to access the raw underlying Python dictionary directly when possible. - Impact: Improved loading performance by ~25%.

3. Loop Inlining (Experiments 6 & 7)

Generic sequence handlers (for List and Dict) are a common bottleneck in Python. - The Optimization: For common types (like List[int] or List[str]), the compiler now inlines the loop directly into the generated handler, avoiding thousands of function calls. - Impact: Reduced loading time for large collections by 35%.

4. Modular Thread-Safety (Experiment 9)

As Lodum moved toward v0.3.0, we needed to support high-concurrency environments. - The Challenge: Ensuring thread-safe access to the global registry and caches without introducing locking overhead. - The Solution: A lock-free fast path for handler lookups using thread-local contexts. - Result: Successfully achieved full thread-safety with zero performance regression.


Key Milestone: The AST Transition

The move from string-based exec() to structured AST (Abstract Syntax Tree) generation was the turning point for the project. By generating and compiling an AST, Lodum: 1. Tailors the Bytecode: The generated code contains the absolute minimum number of operations required for your specific class schema. 2. Improves Debuggability: Even though the code is generated at runtime, the compiler attaches correct line numbers and context, allowing for detailed error reporting without performance penalties. 3. Ensures Safety: Type validation and circular reference checks are baked into the compiled handler, providing native-level safety at Python speed.


For technical details on how we measure these gains, see Benchmark Methodology.