Hbase2.4.0安装

说明:安装zookeeper、hadoop集群

 

解压:

tar -zxvf hbase-2.4.0-bin.tar.gz -C /usr/local

 

改名

cd /usr/local

mv hbase-2.4.0 hbase

 

更改所有者

sudo chown -R hadoop:hadoop hbase

 

查看版本

cd /usr/local/hbase/bin

./hbase version

 

修改profile

vim /etc/profile

export HBASE_HOME=/usr/local/hbase

export PATH=$HBASE_HOME/bin:$PATH

source /etc/profile

 

修改hbase-env.sh

cd /usr/local/hbase/conf/

vim hbase-env.sh

export JAVA_HOME=/usr/lib/jvm/jdk1.8.0_271

export HADOOP_HOME=/usr/local/hdoop

export HBASE_HOME=/usr/local/hbase

# 指定HBase是否使用HBase本身自带的Zookeeper

export HBASE_MANAGES_ZK=false

export HBASE_CLASSPATH=/usr/local/hbase/conf

export PATH=$JAVA_HOME/bin:$HADOOP_HOME/bin:$HBASE_HOME/bin:$PATH

 

修改hbase-site.xml

cd /usr/local/hbase/conf/

vim hbase-site.xml

<configuration>

  <!-- 是否分布式部署 -->

  <property>

      <name>hbase.cluster.distributed</name>

      <value>true</value>

  </property>

     <!-- 指定hbase存放数据的HDFS目录,如果是分布式部署,要和Hadoop的core-site.xml中的fs.defaultFS一致-->

  <property>

      <name>hbase.rootdir</name>

      <value>hdfs://cancer/hbase</value>

      <!—单机模式配置如下

<value>file:///usr/local/hbase/hbase-tmp</value>-->

  </property>

<!-- 配置Zookeeper节点-->

  <property>

      <name>hbase.zookeeper.quorum</name>

      <value>cancer01,cancer02,cancer03,cancer04,cancer05</value>

  </property>

<!-- Hbase在zk上注册的数据信息,默认是/tmp,如果不修改,当系统重启的时候会删除/tmp目录 -->

  <property>

      <name>hbase.zookeeper.property.dataDir</name>

      <value>/usr/local/hbase/zkdata</value>

  </property>

<!-- 设置zk集群端口,默认是2181,一定要和你的zk集群端口保持一致-->

<property>

      <name>hbase.zookeeper.property.clientPort</name>

      <value>2181</value>

</property>

  <property>

      <name>hbase.tmp.dir</name>

      <value>/usr/local/hbase/tmp</value>

  </property>

  <property>

      <name>hbase.unsafe.stream.capability.enforce</name>

      <value>false</value>

  </property>

<!-- HMaster

<property>

       <name>hbase.master</name>

       <value>hdfs://cancer01</value>

  </property>

  <property>

       <name>hbase.wal.provider</name>

       <value>filesystem</value>

  </property>-->

</configuration>

 

修改reginservers

cd /usr/local/hbase/conf/

vim reginservers

#cancer01作为hbase的主节点,部署HMaster,cancer02 03 04 05作为hbase从节点,部署HRegionServer

#此处在cancer01上部署HRegionServer,否则不需要cancer01

cancer01

cancer02

cancer03

cancer04

cancer05

 

拷贝core-site.xml和hdfs-site.xml

为了让Hbase读取到hadoop的配置将两个文件拷贝到 $HBASE_HOME/conf/ 目录下

cp $HADOOP_HOME/etc/hadoop/core-site.xml $HBASE_HOME/conf/

cp $HADOOP_HOME/etc/hadoop/hdfs-site.xml $HBASE_HOME/conf/

 

配置lib

cd /usr/local/hbase/

cp lib/client-facing-thirdparty/htrace-core-3.1.0-incubating.jar lib/

 

复制其他节点

scp -r /usr/local/hbase  hadoop@cancer02:/usr/local/

scp -r /usr/local/hbase  hadoop@cancer03:/usr/local/

scp -r /usr/local/hbase  hadoop@cancer04:/usr/local/

scp -r /usr/local/hbase  hadoop@cancer05:/usr/local/

 

启动zookeeper

zkServer.sh start

 

启动hadoop

cd /usr/local/Hadoop/sbin

./start-all.sh

 

启动hbase

在哪一台机器上输入此 start-hbase.sh启动命令,则那一台机器就是HMaster

cd /usr/local/hbase/bin

./start-hbase.sh

 

在已经有一个HMaster情况下,在一台机器上再启动一个master,该master为备份节点,备份节点可以启动多个(一般启动两个就够了),形成真正的高可用

启动HMaster备份节点,在另外一台机器上启动

hbase-daemon.sh start master

 

启动hbase进程

hbase-daemon.sh start master            -- HMaster

hbase-daemon.sh start regionserver       -- HRegionServer

 

验证

jps

 

 

进入hbase命令行

hbase shell 

status

list

 

web访问

HMaster默认的Web UI端口为16010

http://cancer01:16010/

http://cancer01:16011/master-status

RegionServer默认的Web UI端口为16030

 

在HDFS目录下查看是否生成HBase的数据目录

 

报错现象:用 hbase shell 进入hbase后输入list_namespace,报错Master 正在初始化ERROR:org.apache.hadoop.hbase.PleaseHoldException:Master is initializing

解决方案:

1.进入zk,zkCli.sh -server deptest14:2181

2.删掉meta-region-server目录,rmr /hbase/meta-region-server

如果第一次没有启动成功,第二次尝试启动时必须删掉 hbase.zookeeper.property.dataDir 对应的目录(三台zk上都要操作)

 

Hbase使用

#使用命令进入HBase Shell
$ hbase shell
hbase(main):003:0> 
# HBase提供了大量的帮助文档,只要在HBase 下使用命令help就能够查看HBase所有关键字的帮助
          hbase(main):003:0> help
#如果不知道某个关键字如何使用的话,只需要在Hbase下直接建入该关键字即可
          hbase(main):004:0> put

#在HBase 下,如果输入内容错了,使用回退键是不管用的,必须使用Ctrl+回退键才行

 

命名空间管理

list_namespace
create_namespace 'testdata'
create 'testdata:hb_staff','info'
hbase(main):004:0> put 'testdata:hb_staff','188','info:name','aaron'
hbase(main):005:0> put 'testdata:hb_staff','188','info:age','100'
hbase(main):006:0> put 'testdata:hb_staff','188','info:sex','male'
hbase(main):007:0> put 'testdata:hb_staff','219','info:name','yaoyao'
hbase(main):008:0> put 'testdata:hb_staff','219','info:sex','18'
scan 'testdata:hb_staff'
ROW                                                   COLUMN+CELL                                                                                                                                                
 188                                                  column=info:age, timestamp=1547959759887, value=100                                                                                                        
 188                                                  column=info:name, timestamp=1547959737122, value=aaron                                                                                                     
 188                                                  column=info:sex, timestamp=1547959784771, value=male                                                                                                       
 219                                                  column=info:name, timestamp=1547959820155, value=yaoyao                                                                                                    
 219                                                  column=info:sex, timestamp=1547959831162, value=18                                                                                                         
get 'testdata:hb_staff','188'
COLUMN                                                CELL                                                                                                                                                       
 info:age                                             timestamp=1547959759887, value=100                                                                                                                         
 info:name                                            timestamp=1547959737122, value=aaron                                                                                                                       
 info:sex                                             timestamp=1547959784771, value=male                                                                                                                        
表的管理
1)查看有哪些表
    hbase(main)> list
2)创建表
    # 语法:create <table>, {NAME => <family>, VERSIONS => <VERSIONS>}
    # 例如:创建表t1,有两个family name:f1,f2,且版本数均为2
    hbase(main)> create 't1',{NAME => 'f1', VERSIONS => 2},{NAME => 'f2', VERSIONS => 2}          
        # 创建“student”表,属性有:Sname,Ssex,Sage,Sdept,course
        create 'student','Sname','Ssex','Sage','Sdept','course'        
3)删除表
    分两步:首先disable,然后drop
    例如:删除表t1
    hbase(main)> disable 't1'
    hbase(main)> drop 't1'
4)查看表的结构
    # 语法:describe <table>
    # 例如:查看表t1的结构
    hbase(main)> describe 't1'
5)修改表结构
    修改表结构必须先disable
    # 语法:alter 't1', {NAME => 'f1'}, {NAME => 'f2', METHOD => 'delete'}
    # 例如:修改表test1的cf的TTL为180天
    hbase(main)> disable 'test1'
    hbase(main)> alter 'test1',{NAME=>'body',TTL=>'15552000'},{NAME=>'meta', TTL=>'15552000'}
    hbase(main)> enable 'test1'
权限管理
1)分配权限
# 语法 : grant <user> <permissions> <table> <column family> <column qualifier> 参数后面用逗号分隔
# 权限用五个字母表示: "RWXCA".
# READ('R'), WRITE('W'), EXEC('X'), CREATE('C'), ADMIN('A')
# 例如,给用户‘test'分配对表t1有读写的权限,
hbase(main)> grant 'test','RW','t1'
2)查看权限
    # 语法:user_permission <table>
    # 例如,查看表t1的权限列表
    hbase(main)> user_permission 't1'
3)收回权限
    # 与分配权限类似,语法:revoke <user> <table> <column family> <column qualifier>
    # 例如,收回test用户在表t1上的权限
    hbase(main)> revoke 'test','t1'
表数据的增删改查
1)添加数据
    # 语法:put <table>,<rowkey>,<family:column>,<value>,<timestamp>
    # 例如:给表t1的添加一行记录:rowkey是rowkey001,family name:f1,column name:col1,value:value01,timestamp:系统默认
    hbase(main)> put 't1','rowkey001','f1:col1','value01'
    #用法比较单一。HBase中用put命令添加数据,注意:一次只能为一个表的一行数据的一个列,也就是一个单元格添加一个数据,所以直接用shell命令插入数据效率很低,在实际应用中,一般都是利用编程操作数据。
        put 'student','95001','Sname','LiYing'
        # 为student表添加了学号为95001,名字为LiYing的一行数据,其行键为95001。
        put 'student','95001','course:math','80'   
2)查询数据
    # 1. get命令,用于查看表的某一行数据;2. scan命令用于查看某个表的全部数据
  a)查询某行记录
    # 语法:get <table>,<rowkey>,[<family:column>,....]
    # 例如:查询表t1,rowkey001中的f1下的col1的值
    hbase(main)> get 't1','rowkey001', 'f1:col1'
    # 或者:
    hbase(main)> get 't1','rowkey001', {COLUMN=>'f1:col1'}
    # 查询表t1,rowke002中的f1下的所有列值
    hbase(main)> get 't1','rowkey001'
  b)扫描表
    # 语法:scan <table>, {COLUMNS => [ <family:column>,.... ], LIMIT => num}
    # 另外,还可以添加STARTROW、TIMERANGE和FITLER等高级功能
    # 例如:扫描表t1的前5条数据
    hbase(main)> scan 't1',{LIMIT=>5}
  c)查询表中的数据行数
    # 语法:count <table>, {INTERVAL => intervalNum, CACHE => cacheNum}
    # INTERVAL设置多少行显示一次及对应的rowkey,默认1000;CACHE每次去取的缓存区大小,默认是10,调整该参数可提高查询速度
    # 例如,查询表t1中的行数,每100条显示一次,缓存区为500
    hbase(main)> count 't1', {INTERVAL => 100, CACHE => 500}
3)删除数据
    # 1. delete用于删除一个数据,是put的反向操作;2. deleteall操作用于删除一行数据。
  a )删除行中的某个列值
    # 语法:delete <table>, <rowkey>,  <family:column> , <timestamp>,必须指定列名
    # 例如:删除表t1,rowkey001中的f1:col1的数据
    hbase(main)> delete 't1','rowkey001','f1:col1'
    注:将删除改行f1:col1列所有版本的数据
  b )删除行
    # 语法:deleteall <table>, <rowkey>,  <family:column> , <timestamp>,可以不指定列名,删除整行数据
    # 例如:删除表t1,rowk001的数据
    hbase(main)> deleteall 't1','rowkey001'
  c)删除表中的所有数据
    # 语法: truncate <table>
    # 其具体过程是:disable table -> drop table -> create table
    # 例如:删除表t1的所有数据
    hbase(main)> truncate 't1'
 
4)查询表历史数据
#查询表的历史版本,需要两步。

1、在创建表的时候,指定保存的版本数(假设指定为5)
     create 'teacher',{NAME=>'username',VERSIONS=>5}
2、插入数据然后更新数据,使其产生历史版本数据,注意:这里插入数据和更新数据都是用put命令
put 'teacher','91001','username','Mary'
put 'teacher','91001','username','Mary1'
put 'teacher','91001','username','Mary2'
put 'teacher','91001','username','Mary3'
put 'teacher','91001','username','Mary4'  
put 'teacher','91001','username','Mary5'
3、查询时,指定查询的历史版本数。默认会查询出最新的数据。(有效取值为1到5)
     get 'teacher','91001',{COLUMN=>'username',VERSIONS=>5}
查询结果截图如下:

Region管理
1)移动region
    # 语法:move 'encodeRegionName', 'ServerName'
    # encodeRegionName指的regioName后面的编码,ServerName指的是master-status的Region Servers列表
    # 示例
    hbase(main)>move '4343995a58be8e5bbc739af1e91cd72d', 'db-41.xxx.xxx.org,60020,1390274516739'
2)开启/关闭region
    # 语法:balance_switch true|false
    hbase(main)> balance_switch
3)手动split
    # 语法:split 'regionName', 'splitKey'
4)手动触发major compaction
    #语法:
    #Compact all regions in a table:
    #hbase> major_compact 't1'
    #Compact an entire region:
    #hbase> major_compact 'r1'
    #Compact a single column family within a region:
    #hbase> major_compact 'r1', 'c1'
    #Compact a single column family within a table:
    #hbase> major_compact 't1', 'c1'
配置管理及节点重启
1)修改hdfs配置
    #hdfs配置位置:/etc/hadoop/conf
    # 同步hdfs配置
    cat /home/hadoop/slaves|xargs -i -t scp /etc/hadoop/conf/hdfs-site.xml hadoop@{}:/etc/hadoop/conf/hdfs-site.xml
    #关闭:
    cat /home/hadoop/slaves|xargs -i -t ssh hadoop@{} "sudo /home/hadoop/cdh4/hadoop-2.0.0-cdh4.2.1/sbin/hadoop-daemon.sh --config /etc/hadoop/conf stop datanode"
    #启动:
    cat /home/hadoop/slaves|xargs -i -t ssh hadoop@{} "sudo /home/hadoop/cdh4/hadoop-2.0.0-cdh4.2.1/sbin/hadoop-daemon.sh --config /etc/hadoop/conf start datanode"
2)修改hbase配置
    #hbase配置位置:
    # 同步hbase配置
    cat /home/hadoop/hbase/conf/regionservers|xargs -i -t scp /home/hadoop/hbase/conf/hbase-site.xml hadoop@{}:/home/hadoop/hbase/conf/hbase-site.xml
    # graceful重启
    cd ~/hbase
    bin/graceful_stop.sh --restart --reload --debug inspurXXX.xxx.xxx.org

 

Hbase Java编程

在IDE中添加/usr/local/hbase/lib目录下的jar

新建ExampleForHBase.java,示例简单增删改查,代码如下:

import org.apache.hadoop.conf.Configuration;

import org.apache.hadoop.hbase.*;

import org.apache.hadoop.hbase.client.*;

import org.apache.hadoop.hbase.util.Bytes;

import java.io.IOException;

 

public class ExampleForHBase {

    public static Configuration configuration;

    public static Connection connection;

    public static Admin admin;

    public static void main(String[] args)throws IOException{

        init();

        createTable("student",new String[]{"score"});

        insertData("student","zhangsan","score","English","69");

        insertData("student","zhangsan","score","Math","86");

        insertData("student","zhangsan","score","Computer","77");

        getData("student", "zhangsan", "score","English");

        close();

    }

    public static void init(){

        configuration  = HBaseConfiguration.create();

        configuration.set("hbase.rootdir","hdfs://localhost:9000/hbase");

        try{

            connection = ConnectionFactory.createConnection(configuration);

            admin = connection.getAdmin();

        }catch (IOException e){

            e.printStackTrace();

        }

    }

    public static void close(){

        try{

            if(admin != null){

                admin.close();

            }

            if(null != connection){

                connection.close();

            }

        }catch (IOException e){

            e.printStackTrace();

        }

    }

    public static void createTable(String myTableName,String[] colFamily) throws IOException {

        TableName tableName = TableName.valueOf(myTableName);

        if(admin.tableExists(tableName)){

            System.out.println("talbe is exists!");

        }else {

            TableDescriptorBuilder tableDescriptor = TableDescriptorBuilder.newBuilder(tableName);

            for(String str:colFamily){

                ColumnFamilyDescriptor family =

ColumnFamilyDescriptorBuilder.newBuilder(Bytes.toBytes(str)).build();

                tableDescriptor.setColumnFamily(family);

            }

            admin.createTable(tableDescriptor.build());

        }

    }

    public static void insertData(String tableName,String rowKey,String colFamily,String col,String val) throws IOException {

        Table table = connection.getTable(TableName.valueOf(tableName));

        Put put = new Put(rowKey.getBytes());

        put.addColumn(colFamily.getBytes(),col.getBytes(), val.getBytes());

        table.put(put);

        table.close();

    }

    public static void getData(String tableName,String rowKey,String colFamily, String col)throws  IOException{

        Table table = connection.getTable(TableName.valueOf(tableName));

        Get get = new Get(rowKey.getBytes());

        get.addColumn(colFamily.getBytes(),col.getBytes());

        Result result = table.get(get);

        System.out.println(new String(result.getValue(colFamily.getBytes(),col==null?null:col.getBytes())));

        table.close();

    }

}

 

Hbase行统计的四种方式的效率比对

一、hbase-shell的count命令

行数为 3000W 的表测试结果:

      hbase(main):001:0>count'sda_crm_calls20180102'

 
默认INTERVAL为1000行时花了80分钟。

      hbase(main):001:0>count'sda_crm_calls20180102',INTERVAL=>1000000

 

INTERVAL为1000000行时花了130分钟。

 

二、scan方式设置过滤器循环计数

通过添加 FirstKeyOnlyFilter过滤器的scan进行全表扫描,循环计数RowCount, 速度较慢!但快于第一种count方式!

public void rowCountByScanFilter(String tablename){

    long rowCount = 0;

    try {

        //计时

        StopWatch stopWatch = new StopWatch();

        stopWatch.start();

        TableName name=TableName.valueOf(tablename);

        //connection为类静态变量

        Table table = connection.getTable(name);

        Scan scan = new Scan();

        //FirstKeyOnlyFilter只会取得每行数据的第一个kv,提高count速度

        scan.setFilter(new FirstKeyOnlyFilter());

       

        ResultScanner rs = table.getScanner(scan);

        for (Result result : rs) {

            rowCount += result.size();

        }

        stopWatch.stop();

        System.out.println("RowCount: " + rowCount);

        System.out.println("统计耗时:" +stopWatch.getTotalTimeMillis());

    } catch (Throwable e) {

        e.printStackTrace();

    }

}

 

耗时45分钟!

 

三、利用hbase.RowCounter包执行MR任务

这种方式 效率非常高!利用了hbase jar中自带的统计行数的工具类!

通过 $HBASE_HOME/bin/hbase命令执行:

      [root@cdh1~]# hbase org.apache.hadoop.hbase.mapreduce.RowCounter'sda_crm_calls20180102'

 

耗时 1m40s,速度较上面两种有了质的飞跃!

 

四、利用HBase协处理器Coprocessor(JAVA实现)

协处理器 允许用户在region服务器上运行自己的代码,更准确地说是 允许用户执行region级的操作,并且可以使用与RDBMS中触发器(trigger)类似的功能。。

public void rowCountByCoprocessor(String tablename){

    try {

        //提前创建connection和conf

        Admin admin = connection.getAdmin();

        TableName name=TableName.valueOf(tablename);

        //先disable表,添加协处理器后再enable表

        admin.disableTable(name);

        HTableDescriptor descriptor = admin.getTableDescriptor(name);

        String coprocessorClass = "org.apache.hadoop.hbase.coprocessor.AggregateImplementation";

        if (! descriptor.hasCoprocessor(coprocessorClass)) {

            descriptor.addCoprocessor(coprocessorClass);

        }

        admin.modifyTable(name, descriptor);

        admin.enableTable(name);

        //计时

        StopWatch stopWatch = new StopWatch();

        stopWatch.start();

        Scan scan = new Scan();

        AggregationClient aggregationClient = new AggregationClient(conf);

        System.out.println("RowCount: " + aggregationClient.rowCount(name, new LongColumnInterpreter(), scan));

        stopWatch.stop();

        System.out.println("统计耗时:" +stopWatch.getTotalTimeMillis());

    } catch (Throwable e) {

        e.printStackTrace();

    }

}

 
发现只花了 23秒 就统计完成!

为什么利用协处理器后速度会如此之快?

Table注册了Coprocessor之后,在执行AggregationClient的时候,会将RowCount分散到Table的每一个Region上,Region内RowCount的计算,是通过RPC执行调用接口,由Region对应的RegionServer执行InternalScanner进行的。

因此,性能的提升有两点原因:

1. 分布式统计。将原来客户端按照Rowkey的范围单点进行扫描,然后统计的方式,换成了由所有Region 所在RegionServer同时计算的过程。

2.使用了在RegionServer内部执行使用了 InternalScanner。这是距离实际存储最近的Scanner接口,存取更加快捷。

 

 

更多推荐