巨量資料概論 (Introduction to Big Data) 隨著物聯網、雲端運算與人工智慧的蓬勃發展,巨量資料(Big Data)已成為現代企業數位轉型與決策的核心資產 。本課程專為進修學士班學生設計,以「技術即戰力」為導向,旨在培養具備實作能力的大數據工程與分析人才 。課程內容打破純理論的枯燥框架,採取「基礎觀念 30% + 程式動手實作 50% + 期末專題 20%」的黃金比例,著重工具的親自操作與程式碼撰寫 。課程內容由淺入深,涵蓋四大核心階段:首先從 Linux 環境操作出發,紮實建立 Hadoop HDFS 分散式儲存與 MapReduce 的核心邏輯 ;接著深入學習業界最常用的 SQL-on-Hadoop 工具——Apache Hive,透徹掌握企業級資料建模、內部表與外部表的選型抉擇,以及資料夾分區優化 ;隨後進階至當代計算主流 Apache Spark (PySpark),學習高效能記憶體運算與即時資料流(Kafka、NoSQL)的架構整合 ;最後,銜接雲端大數據中台(Google BigQuery),並結合商業智慧(BI)工具進行資料視覺化 。本課程極具實務價值,評分包含單元上機作業與期中實作檢定,期末更要求學生自由選定公開數據集,完成一套「資料匯入 $\rightarrow$ 資料清洗 $\rightarrow$ 畫布分析」的完整專題作品 。不論未來的志向是成為大數據工程師、ETL 工程師還是資料分析師,本課程都將為您奠定最紮實的履歷亮點與求職即戰力 。
《 課程簡介 -- English 》
Introduction to Big DataWith the rapid growth of the Internet of Things (IoT), cloud computing, and artificial intelligence, Big Data has become a core asset for digital transformation and strategic decision-making in modern enterprises. Designed specifically for undergraduate students in the continuing education program, this course is highly career-oriented, aiming to cultivate practical skills for future big data engineers and data analysts. Moving away from dry, purely theoretical lectures, the curriculum adopts a well-balanced formula: "30% Core Concepts + 50% Hands-on Programming + 20% Final Project," with a heavy emphasis on tool proficiency and coding skills.The curriculum is structured into four progressive phases. It begins with essential Linux command-line operations and builds a solid foundation in distributed storage via Hadoop HDFS and the core logic of MapReduce. Next, students will master Apache Hive, the industry-standard SQL-on-Hadoop tool, to learn enterprise-level data warehousing, the strategic choice between Managed and External tables, and performance optimization via partitioning. The course then advances to Apache Spark (PySpark) for high-performance in-memory computing, alongside stream processing concepts with Apache Kafka and NoSQL databases. Finally, students will explore cloud data platforms like Google BigQuery and connect their analytical outputs to Business Intelligence (BI) tools for interactive data visualization.To ensure hands-on mastery, the grading criteria include lab assignments, a hands-on midterm exam, and a collaborative final project. For the final project, students will select a public dataset to implement a complete pipeline: Data Ingestion $\rightarrow$ ETL/Data Cleaning $\rightarrow$ BI Dashboard Visualization. Whether your career goal is to become a Big Data Engineer, an ETL Engineer, or a Data Analyst, this course will provide you with a powerful portfolio asset and the exact practical skills needed in today's job market.
|