Caching
- Local LLM Inference
—
Caching
,
Machine-Learning
,
Optimization
and +1 more
An empirical systems exploration of local LLM inference on unified memory architectures: profiling compute-bound prefill vs. memory-bound decode phases, analyzing KV cache retention during cancellation, measuring the 64K vs. 128K context cliff on Apple Silicon, and comparing client payload strategies.
- URL Shortener & Pastebin
—
Algorithms
,
Caching
,
Encoding
and +1 more
A robust structural design for a highly available, extremely read-heavy service bridging short aliases to long URLs; implementing Base62 encoding, Snowflake IDs, and strict collision avoidance.
- Distributed Caching Layer for VCS
—
Caching
,
Fault-Tolerance
,
Partitioning
and +2 more
An optimized distributed caching architecture designed to drastically reduce backend I/O and accelerate VCS operations; intelligently caching heavy objects and hashes with ultra-low latency.
- Global CDN Media Serving
—
Caching
,
Edge-Computing
,
Geospatial
and +2 more
A robust edge-optimized CDN-backed media delivery architecture; designed explicitly for seamless, highly available global media serving with completely decoupled background upload processing.
- Proximity Service for Maps
—
Caching
,
Databases
,
Geospatial
and +1 more
An optimized system design for discovering nearby points of interest with ultra-low latency; deeply focusing on efficient spatial indexing via Geohashing, Quadtrees, and read-heavy caching tiers.