Scan with Phone

Scan to instantly open and share this page on your mobile device.

Link copied to clipboard!

Group 9: Streaming & Real-Time Data

Real-time data processing and messaging: Pub/Sub (messaging service), Dataflow (stream processing), Datastream (CDC replication), Apache Kafka on GCE (self-managed streaming). Focus on message ordering, delivery guarantees, latency requirements, and processing patterns.

Stream Processing Hierarchy: Pub/Sub for decoupled messaging; Dataflow for stream transformations; Datastream for database replication; Kafka on GCE for complex streaming architectures. Consider message ordering, exactly-once delivery, and latency needs.

Services & Core Identity

Cloud Pub/Sub

Identity: Fully managed messaging service with global reach and automatic scaling.

Best for: Decoupling microservices, event-driven architectures, fan-out messaging patterns.

Key features: At-least-once delivery, message ordering (when needed), dead letter queues, push/pull subscriptions.

Dataflow

Identity: Serverless stream and batch processing using Apache Beam.

Best for: Real-time analytics, ETL pipelines, stream joins, windowing operations.

Key features: Unified batch/stream model, auto-scaling, exactly-once processing, late data handling.

Datastream

Identity: Change data capture (CDC) service for real-time replication.

Best for: Database synchronization, real-time analytics, migration support.

Key features: Low-latency replication, automatic schema discovery, minimal source impact.

Apache Kafka on GCE

Identity: Self-managed Kafka clusters for complex streaming requirements.

Best for: High-throughput streaming, custom configurations, existing Kafka expertise.

Key features: Full Kafka ecosystem, custom tuning, complete control over configuration.

Key Differences

AspectPub/SubDataflowDatastreamKafka on GCE
ManagementFully managedServerlessFully managedSelf-managed
Primary UseMessagingStream processingDatabase replicationHigh-throughput streaming
OrderingOptional per keyFlexible windowingSource order preservedPer partition
LatencySub-secondSub-second to minutesSub-secondMilliseconds
ScalingAutomaticAutomaticAutomaticManual
ComplexityLowMediumLowHigh

Mathematical Selection Model

Criteria [0..10]. Higher scores indicate better service fit for streaming workloads.

Score_PubSub = 0.35*C_messagingDecoupling + 0.25*C_operationalSimplicity + 0.20*(10 - C_streamProcessing) + 0.15*C_lowLatency + 0.05*C_messageOrdering Score_Dataflow = 0.35*C_streamProcessing + 0.25*C_exactlyOnceProcessing + 0.20*C_operationalSimplicity + 0.15*(10 - C_messagingDecoupling) + 0.05*C_lowLatency Score_Datastream = 0.40*C_databaseReplication + 0.30*C_lowLatency + 0.20*C_operationalSimplicity + 0.10*(10 - C_streamProcessing) Score_KafkaGCE = 0.30*C_highThroughput + 0.25*C_messageOrdering + 0.20*C_lowLatency + 0.15*(10 - C_operationalSimplicity) + 0.10*C_streamProcessing

Current Scores:

{{score.name}}: {{score.value | number:1}}

Interpretation Rules

  • Pub/Sub > 7.0: Ideal for microservices communication, event-driven architectures, and simple messaging patterns.
  • Dataflow > 7.0: Best for complex stream processing, real-time analytics, and ETL pipelines with windowing.
  • Datastream > 7.0: Perfect for database replication, CDC scenarios, and keeping systems in sync.
  • Kafka on GCE > 7.0: Choose for high-throughput streaming, existing Kafka expertise, and complex streaming topologies.

When NOT to Use Streaming Services

  • Batch-only processing: Use BigQuery, Dataproc for large batch analytics without real-time requirements.
  • Simple request-response: Use direct API calls for synchronous communication patterns.
  • File-based ETL: Use Cloud Storage + scheduled jobs for traditional batch processing.
  • Low-volume, infrequent data: Consider simpler polling mechanisms or scheduled tasks.

Summary

Choose streaming services based on your data velocity, processing complexity, and operational preferences. Pub/Sub excels at decoupling, Dataflow at stream processing, Datastream at replication, and Kafka at high-throughput scenarios with complex requirements.

next