日本熟妇hd丰满老熟妇,中文字幕一区二区三区在线不卡 ,亚洲成片在线观看,免费女同在线一区二区

TPC-H Benchmark

云數據庫 SelectDB 版旨在提供卓越的性能和便捷的數據分析服務,在寬表聚合、多表關聯以及高并發點查等場景下均具有優異的性能表現。本文將詳細介紹SelectDB在TPC-H標準測試上的測試方法和測試結果。

概述

TPC-H是一個決策支持基準(Decision Support Benchmark),它由一套面向業務的特別查詢和并發數據修改組成,使用的數據具有廣泛的行業相關性。該基準測試通過一系列的查詢操作來評估數據庫系統在處理復雜查詢和數據挖掘任務時的性能。TPC-H報告的性能指標稱為TPC-H每小時復合查詢性能指標(QphH@Size),該指標反映了系統處理查詢能力的多個方面。這些方面包括執行查詢時所選擇的數據庫大小,由單個流提交查詢時的查詢處理能力,以及由多個并發用戶提交查詢時的查詢吞吐量等。?

說明
  • 本文的TPC-H的實現基于TPC-H的基準測試,并不符合TPC-H基準測試的所有要求。本測試結果不能等同于完全遵守TPC-H測試規范所獲得的測試結果,因此不能與完全遵守該測試規范獲得的測試結果進行對比。

  • TPC-H 在內的標準測試集通常和實際業務場景差距較大,并且部分測試會針對測試集進行參數調優。所以標準測試集的測試結果僅能反映數據庫在特定場景下的性能表現。建議使用實際業務數據進行進一步的測試。

測試環境

  • 數據庫環境。

    環境配置項

    配置說明

    地域和可用區

    華東1(杭州)地域,可用區K。

    規格

    64核 512 GB

    磁盤

    800 GB高性能云硬盤

    云數據庫 SelectDB 版內核版本

    3.0.6

  • 客戶端環境。

    環境配置項

    配置說明

    下載測試工具的設備

    云服務器ECS實例,詳情請參見創建實例。

    地域和可用區

    華東1(杭州)地域

    實例規格

    ecs.g7.2xlarge

    操作系統

    Ubuntu 22.04.1 LTS

    網絡

    云數據庫 SelectDB 版實例為相同專有網絡(VPC)。

測試數據集

整個測試模擬生成TPC-H 100GB和500GB的數據并導入SelectDB進行測試,下面是測試100GB數據表的相關說明及數據量。

TPC-H表名

行數

導入后大小

備注

REGION

5

400KB

區域表

NATION

25

7.714 KB

國家表

SUPPLIER

100萬

85.528 MB

供應商表

PART

2000萬

752.330 MB

零部件表

PARTSUPP

8000萬

4.375 GB

零部件供應表

CUSTOMER

1500萬

1.317 GB

客戶表

ORDERS

1.5億

6.301 GB

訂單表

LINEITEM

6億

20.882 GB

訂單明細表

測試步驟

以下為測試的詳細步驟。如何獲取測試中涉及的腳本,請下載瑤池測試工具

步驟一:安裝unzip工具

  1. 安裝unzip。

    sudo apt install unzip
  2. 安裝完成后,可以通過運行以下命令來驗證unzip是否已成功安裝。

    unzip --help

步驟二:下載安裝TPC-H數據生成工具

從上述腳本庫中獲取腳本后,解壓腳本文件并進入對應目錄,執行以下指令,下載并編譯tpch-dbgen工具,示例如下。

tar -zxvf yaochi_performance_tool.tar.gz
cd ./yaochi_performance_tool/tpch-tools/bin
bash build-tpch-dbgen.sh

安裝成功后,將在TPC-H_Tools_v3.0.0/目錄下生成dbgen二進制文件。

步驟三:生成TPC-H測試集

在安裝測試工具目錄執行以下腳本生成TPC-H數據集,示例如下。

cd ./yaochi_performance_tool/tpch-tools/bin
bash gen-tpch-data.sh

數據會以.tbl為后綴在tpch-data/目錄下生成,默認情況下的文件總大小約100GB。生成時間可能在數分鐘到1小時不等。

步驟四:建表

  1. 準備doris-cluster.conf文件。

    在調用導入腳本前,需要將測試數據庫的連接信息寫在doris-cluster.conf文件中。文件位置在tpch-tools/conf/目錄下,內容包括連接集群的地址、HTTP端口、用戶名、密碼和待導入數據的DB。

    說明

    您可以在云數據庫 SelectDB 版控制臺的實例詳情>網絡信息中獲取VPC地址(或公網地址)和HTTP協議端口。

    export FE_HOST="xxx"
    export FE_HTTP_PORT="8080"
    export FE_QUERY_PORT="9030"
    export USER="root"
    export PASSWORD='xxx'
    export DB="tpch1"
  2. 創建TPC-H表。

    ./yaochi_performance_tool/tpch-tools/bin目錄下執行腳本來自動創建測試用表。

    bash create-tpch-tables.sh

步驟五:導入數據

./yaochi_performance_tool/tpch-tools/bin目錄下執行腳本完成測試集數據導入。

bash ./load-tpch-data.sh

步驟六:檢查導入數據

按照上述流程和參數執行的場合,數據量應和上文測試數據集給出的生成數據的行數一致。

SELECT  COUNT(*) FROM lineitem;
SELECT  COUNT(*) FROM orders;
SELECT  COUNT(*) FROM partsupp;
SELECT  COUNT(*) FROM part;
SELECT  COUNT(*) FROM customer;
SELECT  COUNT(*) FROM supplier;
SELECT  COUNT(*) FROM nation;
SELECT  COUNT(*) FROM region;
SELECT  COUNT(*) FROM revenue0;

步驟七:查詢測試

  • 執行查詢腳本

    執行下面的命令完成查詢測試。

    ./run-tpch-queries.sh

    測試SQL詳情請參見TPCH-Query-SQL。

    說明

    目前SelectDB的查詢優化器和統計信息功能仍有提升空間,所以我們在TPC-H中重寫了一些查詢以適應SelectDB的執行框架,但不影響結果的正確性。

  • 單個 SQL 執行

    下面是本次測試時使用的SQL語句,您也可以從代碼庫中獲取最新的查詢語句,最新測試查詢語句請參見TPC-H 測試查詢語句。

    --ENV config
    SET GLOBAL experimental_enable_nereids_planner=true;
    SET GLOBAL experimental_enable_pipeline_engine=true;
    SET GLOBAL enable_runtime_filter_prune=false;
    SET GLOBAL runtime_filter_wait_time_ms=10000;
    SET GLOBAL enable_fallback_to_original_planner=false;
    SET GLOBAL query_timeout=1000;
    
    --Q1
    SELECT 
        l_returnflag,
        l_linestatus,
        sum(l_quantity) AS sum_qty,
        sum(l_extendedprice) AS sum_base_price,
        sum(l_extendedprice * (1 - l_discount)) AS sum_disc_price,
        sum(l_extendedprice * (1 - l_discount) * (1 + l_tax)) AS sum_charge,
        avg(l_quantity) AS avg_qty,
        avg(l_extendedprice) AS avg_price,
        avg(l_discount) AS avg_disc,
        count(*) AS count_order
    FROM
        lineitem
    WHERE
        l_shipdate <= date '1998-12-01' - interval '90' day
    GROUP BY
        l_returnflag,
        l_linestatus
    ORDER BY
        l_returnflag,
        l_linestatus;
    
    --Q2
    SELECT
        s_acctbal,
        s_name,
        n_name,
        p_partkey,
        p_mfgr,
        s_address,
        s_phone,
        s_comment
    FROM
        partsupp join
        (
            SELECT
                ps_partkey AS a_partkey,
                min(ps_supplycost) AS a_min
            FROM
                partsupp,
                part,
                supplier,
                nation,
                region
            WHERE
                p_partkey = ps_partkey
                AND s_suppkey = ps_suppkey
                AND s_nationkey = n_nationkey
                AND n_regionkey = r_regionkey
                AND r_name = 'EUROPE'
                AND p_size = 15
                AND p_type LIKE '%BRASS'
            GROUP BY a_partkey
        ) AONps_partkey = a_partkey AND ps_supplycost=a_min ,
        part,
        supplier,
        nation,
        region
    WHERE
        p_partkey = ps_partkey
        AND s_suppkey = ps_suppkey
        AND p_size = 15
        AND p_type like '%BRASS'
        AND s_nationkey = n_nationkey
        AND n_regionkey = r_regionkey
        AND r_name = 'EUROPE'
    
    ORDER BY
        s_acctbal DESC,
        n_name,
        s_name,
        p_partkey
    limit 100;
    
    --Q3
    SELECT l_orderkey,
        sum(l_extendedprice * (1 - l_discount)) AS revenue,
        o_orderdate,
        o_shippriority
    FROM
        (
            SELECT l_orderkey, l_extendedprice, l_discount, o_orderdate, o_shippriority, o_custkey FROM
            lineitem JOIN orders
            WHERE l_orderkey = o_orderkey
            AND o_orderdate < date '1995-03-15'
            AND l_shipdate > date '1995-03-15'
        ) t1 JOIN customer c 
       ONc.c_custkey = t1.o_custkey
        WHERE c_mktsegment = 'BUILDING'
    GROUP BY
        l_orderkey,
        o_orderdate,
        o_shippriority
    ORDER BY
        revenue DESC,
        o_orderdate
    limit 10;
    
    --Q4
    SELECT o_orderpriority,
        count(*) AS order_count
    FROM
        (
            SELECT
                *
            FROM
                lineitem
            WHERE l_commitdate < l_receiptdate
        ) t1
        RIGHT semi JOIN orders
       ONt1.l_orderkey = o_orderkey
    WHERE
        o_orderdate >= date '1993-07-01'
        AND o_orderdate < date '1993-07-01' + interval '3' month
    GROUP BY
        o_orderpriority
    ORDER BY
        o_orderpriority;
    
    --Q5
    SELECT n_name,
        sum(l_extendedprice * (1 - l_discount)) AS revenue
    FROM
        customer,
        orders,
        lineitem,
        supplier,
        nation,
        region
    WHERE
        c_custkey = o_custkey
        AND l_orderkey = o_orderkey
        AND l_suppkey = s_suppkey
        AND c_nationkey = s_nationkey
        AND s_nationkey = n_nationkey
        AND n_regionkey = r_regionkey
        AND r_name = 'ASIA'
        AND o_orderdate >= date '1994-01-01'
        AND o_orderdate < date '1994-01-01' + interval '1' year
    GROUP BY
        n_name
    ORDER BY
        revenue DESC;
    
    --Q6
    SELECT sum(l_extendedprice * l_discount) AS revenue
    FROM
        lineitem
    WHERE
        l_shipdate >= date '1994-01-01'
        AND l_shipdate < date '1994-01-01' + interval '1' year
        AND l_discount between .06 - 0.01 AND .06 + 0.01
        AND l_quantity < 24;
    
    --Q7
    SELECT supp_nation,
        cust_nation,
        l_year,
        sum(volume) AS revenue
    FROM
        (
            SELECT
                n1.n_name AS supp_nation,
                n2.n_name AS cust_nation,
                extract(year FROM l_shipdate) AS l_year,
                l_extendedprice * (1 - l_discount) AS volume
            FROM
                supplier,
                lineitem,
                orders,
                customer,
                nation n1,
                nation n2
            WHERE
                s_suppkey = l_suppkey
                AND o_orderkey = l_orderkey
                AND c_custkey = o_custkey
                AND s_nationkey = n1.n_nationkey
                AND c_nationkey = n2.n_nationkey
                AND (
                    (n1.n_name = 'FRANCE' AND n2.n_name = 'GERMANY')
                    or (n1.n_name = 'GERMANY' AND n2.n_name = 'FRANCE')
                )
                AND l_shipdate between date '1995-01-01' AND date '1996-12-31'
        ) AS shipping
    GROUP BY
        supp_nation,
        cust_nation,
        l_year
    ORDER BY
        supp_nation,
        cust_nation,
        l_year;
    
    --Q8
    
    SELECT o_year,
        sum(case
            when nation = 'BRAZIL' then volume
            else 0
        end) / sum(volume) AS mkt_share
    FROM
        (
            SELECT
                extract(year FROM o_orderdate) AS o_year,
                l_extendedprice * (1 - l_discount) AS volume,
                n2.n_name AS nation
            FROM
                lineitem,
                orders,
                customer,
                supplier,
                part,
                nation n1,
                nation n2,
                region
            WHERE
                p_partkey = l_partkey
                AND s_suppkey = l_suppkey
                AND l_orderkey = o_orderkey
                AND o_custkey = c_custkey
                AND c_nationkey = n1.n_nationkey
                AND n1.n_regionkey = r_regionkey
                AND r_name = 'AMERICA'
                AND s_nationkey = n2.n_nationkey
                AND o_orderdate between date '1995-01-01' AND date '1996-12-31'
                AND p_type = 'ECONOMY ANODIZED STEEL'
        ) AS all_nations
    GROUP BY
        o_year
    ORDER BY
        o_year;
    
    --Q9
    SELECT nation,
        o_year,
        sum(amount) AS sum_profit
    FROM
        (
            SELECT
                n_name AS nation,
                extract(year FROM o_orderdate) AS o_year,
                l_extendedprice * (1 - l_discount) - ps_supplycost * l_quantity AS amount
            FROM
                lineitem JOIN ordersONo_orderkey = l_orderkey
                join[shuffle] partONp_partkey = l_partkey
                join[shuffle] partsuppONps_partkey = l_partkey
                join[shuffle] supplierONs_suppkey = l_suppkey
                join[broadcast] nationONs_nationkey = n_nationkey
            WHERE
                ps_suppkey = l_suppkey AND 
                p_name like '%green%'
        ) AS profit
    GROUP BY
        nation,
        o_year
    ORDER BY
        nation,
        o_year DESC;
    
    --Q10
    SELECT c_custkey,
        c_name,
        sum(t1.l_extendedprice * (1 - t1.l_discount)) AS revenue,
        c_acctbal,
        n_name,
        c_address,
        c_phone,
        c_comment
    FROM
        customer,
        (
            SELECT o_custkey,l_extendedprice,l_discount FROM lineitem, orders
            WHERE l_orderkey = o_orderkey
            AND o_orderdate >= date '1993-10-01'
            AND o_orderdate < date '1993-10-01' + interval '3' month
            AND l_returnflag = 'R'
        ) t1,
        nation
    WHERE
        c_custkey = t1.o_custkey
        AND c_nationkey = n_nationkey
    GROUP BY
        c_custkey,
        c_name,
        c_acctbal,
        c_phone,
        n_name,
        c_address,
        c_comment
    ORDER BY
        revenue DESC
    limit 20;
    
    --Q11
    SELECT ps_partkey,
        sum(ps_supplycost * ps_availqty) AS value
    FROM
        partsupp,
        (
        SELECT s_suppkey
        FROM supplier, nation
        WHERE s_nationkey = n_nationkey AND n_name = 'GERMANY'
        ) B
    WHERE
        ps_suppkey = B.s_suppkey
    GROUP BY
        ps_partkey having
            sum(ps_supplycost * ps_availqty) > (
                SELECT
                    sum(ps_supplycost * ps_availqty) * 0.000002
                FROM
                    partsupp,
                    (SELECT s_suppkey
                     FROM supplier, nation
                     WHERE s_nationkey = n_nationkey AND n_name = 'GERMANY'
                    ) A
                WHERE
                    ps_suppkey = A.s_suppkey
            )
    ORDER BY
        value DESC;
    
    --Q12
    SELECT l_shipmode,
        sum(case
            when o_orderpriority = '1-URGENT'
                or o_orderpriority = '2-HIGH'
                then 1
            else 0
        end) AS high_line_count,
        sum(case
            when o_orderpriority <> '1-URGENT'
                AND o_orderpriority <> '2-HIGH'
                then 1
            else 0
        end) AS low_line_count
    FROM
        orders,
        lineitem
    WHERE
        o_orderkey = l_orderkey
        AND l_shipmode in ('MAIL', 'SHIP')
        AND l_commitdate < l_receiptdate
        AND l_shipdate < l_commitdate
        AND l_receiptdate >= date '1994-01-01'
        AND l_receiptdate < date '1994-01-01' + interval '1' year
    GROUP BY
        l_shipmode
    ORDER BY
        l_shipmode;
    
    --Q13
    SELECT c_count,
        count(*) AS custdist
    FROM
        (
            SELECT
                c_custkey,
                count(o_orderkey) AS c_count
            FROM
                orders RIGHT outer JOIN customer on
                    c_custkey = o_custkey
                    AND o_comment not like '%special%requests%'
            GROUP BY
                c_custkey
        ) AS c_orders
    GROUP BY
        c_count
    ORDER BY
        custdist DESC,
        c_count DESC;
    
    --Q14
    SELECT 100.00 * sum(case
            when p_type like 'PROMO%'
                then l_extendedprice * (1 - l_discount)
            else 0
        end) / sum(l_extendedprice * (1 - l_discount)) AS promo_revenue
    FROM
        part,
        lineitem
    WHERE
        l_partkey = p_partkey
        AND l_shipdate >= date '1995-09-01'
        AND l_shipdate < date '1995-09-01' + interval '1' month;
    
    --Q15
    SELECT s_suppkey,
        s_name,
        s_address,
        s_phone,
        total_revenue
    FROM
        supplier,
        revenue0
    WHERE
        s_suppkey = supplier_no
        AND total_revenue = (
            SELECT
                max(total_revenue)
            FROM
                revenue0
        )
    ORDER BY
        s_suppkey;
    
    --Q16
    SELECT p_brAND,
        p_type,
        p_size,
        count(distinct ps_suppkey) AS supplier_cnt
    FROM
        partsupp,
        part
    WHERE
        p_partkey = ps_partkey
        AND p_brAND <> 'BrAND#45'
        AND p_type not like 'MEDIUM POLISHED%'
        AND p_size in (49, 14, 23, 45, 19, 3, 36, 9)
        AND ps_suppkey not in (
            SELECT
                s_suppkey
            FROM
                supplier
            WHERE
                s_comment like '%Customer%Complaints%'
        )
    GROUP BY
        p_brAND,
        p_type,
        p_size
    ORDER BY
        supplier_cnt DESC,
        p_brAND,
        p_type,
        p_size;
    
    --Q17
    SELECT sum(l_extendedprice) / 7.0 AS avg_yearly
    FROM
        lineitem JOIN [broadcast]
        part p1ONp1.p_partkey = l_partkey
    WHERE
        p1.p_brAND = 'BrAND#23'
        AND p1.p_container = 'MED BOX'
        AND l_quantity < (
            SELECT
                0.2 * avg(l_quantity)
            FROM
                lineitem JOIN [broadcast]
                part p2ONp2.p_partkey = l_partkey
            WHERE
                l_partkey = p1.p_partkey
                AND p2.p_brAND = 'BrAND#23'
                AND p2.p_container = 'MED BOX'
        );
    
    --Q18
    SELECT c_name,
        c_custkey,
        t3.o_orderkey,
        t3.o_orderdate,
        t3.o_totalprice,
        sum(t3.l_quantity)
    FROM
    customer join
    (
      SELECT * FROM
      lineitem join
      (
        SELECT * FROM
        orders left semi join
        (
          SELECT
              l_orderkey
          FROM
              lineitem
          GROUP BY
              l_orderkey having sum(l_quantity) > 300
        ) t1
       ONo_orderkey = t1.l_orderkey
      ) t2
     ONt2.o_orderkey = l_orderkey
    ) t3
    on c_custkey = t3.o_custkey
    GROUP BY
        c_name,
        c_custkey,
        t3.o_orderkey,
        t3.o_orderdate,
        t3.o_totalprice
    ORDER BY
        t3.o_totalprice DESC,
        t3.o_orderdate
    limit 100;
    
    --Q19
    SELECT sum(l_extendedprice* (1 - l_discount)) AS revenue
    FROM
        lineitem,
        part
    WHERE
        (
            p_partkey = l_partkey
            AND p_brAND = 'BrAND#12'
            AND p_container in ('SM CASE', 'SM BOX', 'SM PACK', 'SM PKG')
            AND l_quantity >= 1 AND l_quantity <= 1 + 10
            AND p_size between 1 AND 5
            AND l_shipmode in ('AIR', 'AIR REG')
            AND l_shipinstruct = 'DELIVER IN PERSON'
        )
        or
        (
            p_partkey = l_partkey
            AND p_brAND = 'BrAND#23'
            AND p_container in ('MED BAG', 'MED BOX', 'MED PKG', 'MED PACK')
            AND l_quantity >= 10 AND l_quantity <= 10 + 10
            AND p_size between 1 AND 10
            AND l_shipmode in ('AIR', 'AIR REG')
            AND l_shipinstruct = 'DELIVER IN PERSON'
        )
        or
        (
            p_partkey = l_partkey
            AND p_brAND = 'BrAND#34'
            AND p_container in ('LG CASE', 'LG BOX', 'LG PACK', 'LG PKG')
            AND l_quantity >= 20 AND l_quantity <= 20 + 10
            AND p_size between 1 AND 15
            AND l_shipmode in ('AIR', 'AIR REG')
            AND l_shipinstruct = 'DELIVER IN PERSON'
        );
    
    --Q20
    SELECT s_name, s_address FROM
    supplier left semi join
    (
        SELECT * FROM
        (
            SELECT l_partkey,l_suppkey, 0.5 * sum(l_quantity) AS l_q
            FROM lineitem
            WHERE l_shipdate >= date '1994-01-01'
                AND l_shipdate < date '1994-01-01' + interval '1' year
            GROUP BY l_partkey,l_suppkey
        ) t2 join
        (
            SELECT ps_partkey, ps_suppkey, ps_availqty
            FROM partsupp left semi JOIN part
           ONps_partkey = p_partkey AND p_name like 'forest%'
        ) t1
       ONt2.l_partkey = t1.ps_partkey AND t2.l_suppkey = t1.ps_suppkey
        AND t1.ps_availqty > t2.l_q
    ) t3
    on s_suppkey = t3.ps_suppkey
    join nation
    WHERE s_nationkey = n_nationkey
        AND n_name = 'CANADA'
    ORDER BY s_name;
    
    --Q21
    SELECT s_name, count(*) AS numwait
    FROM
      lineitem l2 RIGHT semi join
      (
        SELECT * FROM
        lineitem l3 RIGHT anti join
        (
          SELECT * FROM
          orders JOIN lineitem l1ONl1.l_orderkey = o_orderkey AND o_orderstatus = 'F'
          join
          (
            SELECT * FROM
            supplier JOIN nation
            WHERE s_nationkey = n_nationkey
              AND n_name = 'SAUDI ARABIA'
          ) t1
          WHERE t1.s_suppkey = l1.l_suppkey AND l1.l_receiptdate > l1.l_commitdate
        ) t2
       ONl3.l_orderkey = t2.l_orderkey AND l3.l_suppkey <> t2.l_suppkey  AND l3.l_receiptdate > l3.l_commitdate
      ) t3
     ONl2.l_orderkey = t3.l_orderkey AND l2.l_suppkey <> t3.l_suppkey 
    
    GROUP BY
        t3.s_name
    ORDER BY
        numwait DESC,
        t3.s_name
    limit 100;
    
    --Q22
    with tmp AS (SELECT
                        avg(c_acctbal) AS av
                    FROM
                        customer
                    WHERE
                        c_acctbal > 0.00
                        AND substring(c_phone, 1, 2) in
                            ('13', '31', '23', '29', '30', '18', '17'))
    
    SELECT cntrycode,
        count(*) AS numcust,
        sum(c_acctbal) AS totacctbal
    FROM
        (
        SELECT
                substring(c_phone, 1, 2) AS cntrycode,
                c_acctbal
            FROM
                 orders RIGHT anti JOIN customer cON o_custkey = c.c_custkey JOIN tmpONc.c_acctbal > tmp.av
            WHERE
                substring(c_phone, 1, 2) in
                    ('13', '31', '23', '29', '30', '18', '17')
        ) AS custsale
    GROUP BY
        cntrycode
    ORDER BY
        cntrycode;

測試結果

以下為TPCH 100 GB和500 GB的測試結果。

Query

TPCH 100GB(s)

TPCH 500GB(s)

Q1

1.74

10.04

Q2

0.07

0.19

Q3

0.34

3.43

Q4

0.19

1.1

Q5

0.81

7.52

Q6

0.03

0.15

Q7

0.54

5.74

Q8

0.26

2.56

Q9

2.62

18.44

Q10

0.91

5.45

Q11

0.08

0.36

Q12

0.09

0.47

Q13

1.32

6.6

Q14

0.12

0.59

Q15

0.18

0.85

Q16

0.28

1.17

Q17

0.1

0.45

Q18

1.7

9.94

Q19

0.18

1.9

Q20

0.39

0.62

Q21

0.65

7.23

Q22

0.19

1.04

合計

12.79

85.84