Designing a Cloud-Native Lakehouse Architecture for Real-Time and Batch Enterprise Data Intelligence
DOI:
https://doi.org/10.70917/ijcisim-2026-5512Keywords:
batch analytics, cloud-native architecture, data governance, data lakehouse, medallion architecture, real-time ingestion, streaming analytics, transactional table formatsAbstract
Enterprises require unified platforms capable of supporting real-time data ingestion and large-scale batch analytics within the same operational environment. Traditional architectures separate data lakes, data warehouses, and streaming systems, producing duplicated storage, fragmented governance, inconsistent data semantics, and escalating infrastructure costs. This article proposes a cloud-native lakehouse architecture that converges real-time and batch data processing within a single scalable framework. The architecture integrates six functional layers: ingestion, storage with transactional table formats, distributed processing, semantic serving, governance, and workload optimization. A medallion data organization model structures the progression of data through three quality zones, raw ingestion, conformed and validated datasets, and curated business-ready outputs, from source to analytical consumption. The proposed architecture is evaluated using metrics including ingestion throughput, end-to-end latency, query performance, storage efficiency, cost per workload, and schema evolution ease. Expected benefits include reduced data duplication, lower operational overhead from unified governance, more consistent data quality across workloads, and improved support for artificial intelligence and machine learning pipelines. The contribution is a reference architecture and implementation blueprint for enterprise-scale data intelligence that unifies streaming and batch processing under a consistent storage, governance, and serving model.