Search Posts
Search blog posts by title, description, tags, or content.
Back to About
SaaS Monitoring Data Collection Pipeline
EXEM · Jan 2021 - Sep 2021
Kafka StreamApache DruidSpring BootKubernetes
Tasks
- •Used Kafka Streams to transform collected Kafka data into service-specific query data
- •Collected, stored, and served real-time data with Apache Druid
- •Optimized Apache Druid ingestion and query paths under limited infrastructure
- •Configured Kubernetes scale in/out based on memory usage for uninterrupted operation under changing load
Achievements
- •Implemented Kafka Streams with Spring Boot for maintainability
- •Reduced duplicated data-processing logic across servers and improved productivity
- •Maintained 45,000 events per second even with about half of the officially recommended infrastructure scale
- •Reduced data loss and duplication to maintain collection consistency