Enterprise Data Warehouse Optimization with Hadoop on IBM Power Systems Servers

Name: Enterprise Data Warehouse Optimization with Hadoop on IBM Power Systems Servers
Rating: 4.333333492279053 (3 reviews)

Scott Vetter · Helen Lu · Maciej Olejniczak · IBM Redbooks

thg 1 2018 · IBM Redbooks

4,3

3 bài đánh giá

Sách điện tử

Trang

Đủ điều kiện

Điểm xếp hạng và bài đánh giá chưa được xác minh Tìm hiểu thêm

Giới thiệu về sách điện tử này

Data warehouses were developed for many good reasons, such as providing quick query and reporting for business operations, and business performance. However, over the years, due to the explosion of applications and data volume, many existing data warehouses have become difficult to manage. Extract, Transform, and Load (ETL) processes are taking longer, missing their allocated batch windows. In addition, data types that are required for business analysis have expanded from structured data to unstructured data.

The Apache open source Hadoop platform provides a great alternative for solving these problems.

IBM® has committed to open source since the early years of open Linux. IBM and Hortonworks together are committed to Apache open source software more than any other company.

IBM Power SystemsTM servers are built with open technologies and are designed for mission-critical data applications. Power Systems servers use technology from the OpenPOWER Foundation, an open technology infrastructure that uses the IBM POWER® architecture to help meet the evolving needs of big data applications. The combination of Power Systems with Hortonworks Data Platform (HDP) provides users with a highly efficient platform that provides leadership performance for big data workloads such as Hadoop and Spark.

This IBM RedpaperTM publication provides details about Enterprise Data Warehouse (EDW) optimization with Hadoop on Power Systems. Many people know Power Systems from the IBM AIX® platform, but might not be familiar with IBM PowerLinuxTM, so part of this paper provides a Power Systems overview. A quick introduction to Hadoop is provided for those not familiar with the topic. Details of HDP on Power Reference architecture are included that will help both software architects and infrastructure architects understand the design.

In the optimization chapter, we describe various topics: traditional EDW offload, sizing guidelines, performance tuning, IBM Elastic StorageTM Server (ESS) for data-intensive workload, IBM Big SQL as the common structured query language (SQL) engine for Hadoop platform, and tools that are available on Power Systems that are related to EDW optimization. We also dedicate some pages to the analytics components (IBM Data Science Experience (IBM DSX) and IBM SpectrumTM Conductor for Spark workload) for the Hadoop infrastructure.

Xếp hạng và đánh giá

4,3

3 bài đánh giá

Xếp hạng sách điện tử này

Cho chúng tôi biết suy nghĩ của bạn.

Đọc thông tin

Điện thoại thông minh và máy tính bảng

Cài đặt ứng dụng Google Play Sách cho Android và iPad/iPhone. Ứng dụng sẽ tự động đồng bộ hóa với tài khoản của bạn và cho phép bạn đọc trực tuyến hoặc ngoại tuyến dù cho bạn ở đâu.

Máy tính xách tay và máy tính

Bạn có thể nghe các sách nói đã mua trên Google Play thông qua trình duyệt web trên máy tính.

Thiết bị đọc sách điện tử và các thiết bị khác

Để đọc trên thiết bị e-ink như máy đọc sách điện tử Kobo, bạn sẽ cần tải tệp xuống và chuyển tệp đó sang thiết bị của mình. Hãy làm theo hướng dẫn chi tiết trong Trung tâm trợ giúp để chuyển tệp sang máy đọc sách điện tử được hỗ trợ.

Báo cáo nội dung bất hợp pháp