Group 9: Streaming & Real-Time Data
Real-time data processing and messaging: Pub/Sub (messaging service), Dataflow (stream processing), Datastream (CDC replication), Apache Kafka on GCE (self-managed streaming). Focus on message ordering, delivery guarantees, latency requirements, and processing patterns.
Services & Core Identity
Cloud Pub/Sub
Identity: Fully managed messaging service with global reach and automatic scaling.
Best for: Decoupling microservices, event-driven architectures, fan-out messaging patterns.
Key features: At-least-once delivery, message ordering (when needed), dead letter queues, push/pull subscriptions.
Dataflow
Identity: Serverless stream and batch processing using Apache Beam.
Best for: Real-time analytics, ETL pipelines, stream joins, windowing operations.
Key features: Unified batch/stream model, auto-scaling, exactly-once processing, late data handling.
Datastream
Identity: Change data capture (CDC) service for real-time replication.
Best for: Database synchronization, real-time analytics, migration support.
Key features: Low-latency replication, automatic schema discovery, minimal source impact.
Apache Kafka on GCE
Identity: Self-managed Kafka clusters for complex streaming requirements.
Best for: High-throughput streaming, custom configurations, existing Kafka expertise.
Key features: Full Kafka ecosystem, custom tuning, complete control over configuration.
Key Differences
| Aspect | Pub/Sub | Dataflow | Datastream | Kafka on GCE |
|---|---|---|---|---|
| Management | Fully managed | Serverless | Fully managed | Self-managed |
| Primary Use | Messaging | Stream processing | Database replication | High-throughput streaming |
| Ordering | Optional per key | Flexible windowing | Source order preserved | Per partition |
| Latency | Sub-second | Sub-second to minutes | Sub-second | Milliseconds |
| Scaling | Automatic | Automatic | Automatic | Manual |
| Complexity | Low | Medium | Low | High |
Mathematical Selection Model
Criteria [0..10]. Higher scores indicate better service fit for streaming workloads.
Current Scores:
Interpretation Rules
- Pub/Sub > 7.0: Ideal for microservices communication, event-driven architectures, and simple messaging patterns.
- Dataflow > 7.0: Best for complex stream processing, real-time analytics, and ETL pipelines with windowing.
- Datastream > 7.0: Perfect for database replication, CDC scenarios, and keeping systems in sync.
- Kafka on GCE > 7.0: Choose for high-throughput streaming, existing Kafka expertise, and complex streaming topologies.
When NOT to Use Streaming Services
- Batch-only processing: Use BigQuery, Dataproc for large batch analytics without real-time requirements.
- Simple request-response: Use direct API calls for synchronous communication patterns.
- File-based ETL: Use Cloud Storage + scheduled jobs for traditional batch processing.
- Low-volume, infrequent data: Consider simpler polling mechanisms or scheduled tasks.
Summary
Choose streaming services based on your data velocity, processing complexity, and operational preferences. Pub/Sub excels at decoupling, Dataflow at stream processing, Datastream at replication, and Kafka at high-throughput scenarios with complex requirements.