Devy

Search Posts

Search blog posts by title, description, tags, or content.

Back to About

SaaS Monitoring Data Collection Pipeline

EXEM · Jan 2021 - Sep 2021

Kafka StreamApache DruidSpring BootKubernetes

Tasks

  • Used Kafka Streams to transform collected Kafka data into service-specific query data
  • Collected, stored, and served real-time data with Apache Druid
  • Optimized Apache Druid ingestion and query paths under limited infrastructure
  • Configured Kubernetes scale in/out based on memory usage for uninterrupted operation under changing load

Achievements

  • Implemented Kafka Streams with Spring Boot for maintainability
  • Reduced duplicated data-processing logic across servers and improved productivity
  • Maintained 45,000 events per second even with about half of the officially recommended infrastructure scale
  • Reduced data loss and duplication to maintain collection consistency