Performance Evolution: The AST Revolution
Lodum's performance is the result of systematic experimentation and a fundamental shift in how Python serialization is handled. This document tracks our journey from a dynamic baseline to a high-performance, bytecode-compiled engine.
Performance Trajectory
Standardized on Linux (ubuntu-latest) hardware. Measurements in microseconds (us) per operation.
| Scenario | v0.1.0 (Initial) | v0.2.0 (AST Core) | v0.3.0 (Modular) | Total Improvement |
|---|---|---|---|---|
| Simple Serialization | 5.56 | 4.86 | 4.90 | 12% faster |
| Simple Deserialization | 18.82 | 12.05 | 12.47 | 34% faster |
| Complex Serialization | 10.95 | 7.79 | 7.90 | 28% faster |
| Complex Deserialization | 37.31 | 25.99 | 26.35 | 29% faster |
| Nested Serialization | 25.58 | 21.54 | 21.69 | 15% faster |
| Nested Deserialization | 117.81 | 72.63 | 74.44 | 37% faster |
Core Optimization Experiments
Our optimization strategy focused on eliminating the "Interpreter Tax"—the overhead of generic loops and repeated runtime inspections.
1. Pre-Resolved Handlers (Experiment 2)
Originally, every field access triggered a cache lookup and type introspection. - The Optimization: We moved handler resolution to compile-time. The generated function now has direct references to the handlers it needs. - Impact: Achieved a 65% reduction in dumping overhead.
2. Fast Raw-Dict Access (Experiment 5)
Traditional loaders often wrap input data in several layers of abstraction. - The Optimization: We implemented a "Fast Path" that allows the generated bytecode to access the raw underlying Python dictionary directly when possible. - Impact: Improved loading performance by ~25%.
3. Loop Inlining (Experiments 6 & 7)
Generic sequence handlers (for List and Dict) are a common bottleneck in Python.
- The Optimization: For common types (like List[int] or List[str]), the compiler now inlines the loop directly into the generated handler, avoiding thousands of function calls.
- Impact: Reduced loading time for large collections by 35%.
4. Modular Thread-Safety (Experiment 9)
As Lodum moved toward v0.3.0, we needed to support high-concurrency environments. - The Challenge: Ensuring thread-safe access to the global registry and caches without introducing locking overhead. - The Solution: A lock-free fast path for handler lookups using thread-local contexts. - Result: Successfully achieved full thread-safety with zero performance regression.
Key Milestone: The AST Transition
The move from string-based exec() to structured AST (Abstract Syntax Tree) generation was the turning point for the project. By generating and compiling an AST, Lodum:
1. Tailors the Bytecode: The generated code contains the absolute minimum number of operations required for your specific class schema.
2. Improves Debuggability: Even though the code is generated at runtime, the compiler attaches correct line numbers and context, allowing for detailed error reporting without performance penalties.
3. Ensures Safety: Type validation and circular reference checks are baked into the compiled handler, providing native-level safety at Python speed.
For technical details on how we measure these gains, see Benchmark Methodology.