This article is about learning how to delete bulk column values by using Hbase bulk loading with Hadoop MapReduce. Proficient hadoop developers are sharing important things required for bulk column deletion in hadoop development. You can follow the steps shared by them to know how they do it.
We are introducing a way to delete multiple column values by using hbasebulkloading feature with hadoopmapreduce.
Problem:
Consider a case we have loaded the peta bytes sized data in an hbase table with billions of rows, now we have hundreds of columns in that data, our interest is to delete the columns and its data which is not required for our purpose.
There are basically 2 solutions to achieve that purpose.
1) Using HBase client API with Java and perform Delete operations.
2) Using HBase’s bulk load feature, scan whole table and delete selected column values from all rows.
From above 2 solutions the first solution is indeed a rudimentary solution because it will take awaful lot of time and it will fire lot of DeleteRequests which is time consuming approach while approach 2 which uses powerful feature of BulkLoading supported by HBase and leveraging the Hadoop MapReduce paradigm will do the megic.
Indeed bulkloading the data by generating HFiles using hadoopmapreduce will bypass the WAL so it will take less time and hence it will be time and resource efficient.
Solution:
Lets consider a simple case for our example, we have a table named user in hbase which is having 2 column families userDetails and personalDetails.
So my data looks like,

so, from above data we are not interest in city, streetAddress and secondaryEmail columns which belongs to personalDetails and userDetails column family respectively.
I am using a property file which will list the column family and column’s mapping which we will use to delete that column values.
My property file looks like,

So, in order to delete the column values below is my solution code.
BulkDeleteColumnValueDriver.java
importjava.io.BufferedReader;
importjava.io.File;
importjava.io.FileReader;
importjava.util.ArrayList;
importjava.util.List;
importorg.apache.hadoop.conf.Configuration;
importorg.apache.hadoop.conf.Configured;
importorg.apache.hadoop.fs.Path;
importorg.apache.hadoop.hbase.HBaseConfiguration;
importorg.apache.hadoop.hbase.KeyValue;
importorg.apache.hadoop.hbase.client.HTable;
importorg.apache.hadoop.hbase.client.Scan;
importorg.apache.hadoop.hbase.io.ImmutableBytesWritable;
importorg.apache.hadoop.hbase.mapreduce.HFileOutputFormat;
importorg.apache.hadoop.hbase.mapreduce.LoadIncrementalHFiles;
importorg.apache.hadoop.hbase.mapreduce.TableMapReduceUtil;
importorg.apache.hadoop.hbase.util.Bytes;
importorg.apache.hadoop.mapreduce.Job;
importorg.apache.hadoop.mapreduce.lib.output.FileOutputFormat;
importorg.apache.hadoop.util.Tool;
importorg.apache.hadoop.util.ToolRunner;
publicclassBulkDeleteColumnValueDriverextends Configured implements Tool
{
privatestaticfinal String PROPERTY_VALUE_SEPERATOR = “~”;
privatestatic List<String>columnFamilyColumnMappingList = newArrayList<String>();
@Override
publicint run(String[] args) throws Exception
{
Configuration configuration = HBaseConfiguration.create();
configuration.set(“mapred.map.tasks.speculative.execution”, “false”);
String tableName = args[0];
HTablehbaseTable = newHTable(configuration, tableName);
/*Reading property file and filling the mapping list*/
initializeColumnMapping(args[1]);
Job job = newJob(configuration, “Bulk Delete Column Values Test“);
job.setJarByClass(BulkDeleteColumnValueDriver.class);
Scan scan = newScan();
scan.setCacheBlocks(false);
scan.setCaching(1000);
/*Iterating, parsing and adding the columnfamily& column to
*scan the added columns will be marked for deletion */
for(String mappingLine : columnFamilyColumnMappingList)
{
String[] values = mappingLine.split(PROPERTY_VALUE_SEPERATOR);
scan.addColumn(Bytes.toBytes(values[0]), Bytes.toBytes(values[1]));
}
TableMapReduceUtil.initTableMapperJob(tableName, scan, BulkDeleteColumnValueMapper.class, ImmutableBytesWritable.class, KeyValue.class, job);
HFileOutputFormat.configureIncrementalLoad(job, hbaseTable);
String outputPath = “/tmp/bulkDeleteColumnValues”;
FileOutputFormat.setOutputPath(job, new Path(outputPath));
job.waitForCompletion(true);
if(job.isSuccessful())
{
LoadIncrementalHFilesloadFfiles = newLoadIncrementalHFiles(configuration);
HTabletable = newHTable(configuration, tableName);
loadFfiles.doBulkLoad(new Path(outputPath), table);
System.out.println(“Successful”);
}
else
{
System.out.println(“Error in Job..”);
}
hbaseTable.flushCommits();
hbaseTable.close();
return 0;
}
privatestaticvoidinitializeColumnMapping(String filePath)
{
try
{
BufferedReaderreader = newBufferedReader(newFileReader(new File(filePath)));
String line;
while((line = reader.readLine()) != null)
{
// Skipping comment lines
if(!line.startsWith(“##”) &&line.length() > 0)
{
if(line.contains(PROPERTY_VALUE_SEPERATOR))
{
columnFamilyColumnMappingList.add(line);
}
}
}
reader.close();
}
catch(Exception exception)
{
exception.printStackTrace();
}
}
publicstaticvoid main(String[] args) throws Exception
{
if(args.length == 2)
{
intexitCode = ToolRunner.run(HBaseConfiguration.create(), newBulkDeleteColumnValueDriver(), args);
System.exit(exitCode);
}
else
{
System.out.println(“Usage:<HBaseTableName><ColumnMappingFilePath>”);
}
}
}
BulkDeleteColumnValueMapper.java
importjava.io.IOException;
importorg.apache.hadoop.hbase.HConstants;
importorg.apache.hadoop.hbase.KeyValue;
importorg.apache.hadoop.hbase.client.Result;
importorg.apache.hadoop.hbase.io.ImmutableBytesWritable;
importorg.apache.hadoop.hbase.mapreduce.TableMapper;
publicclassBulkDeleteColumnValueMapperextendsTableMapper<ImmutableBytesWritable, KeyValue>
{
@Override
publicvoid map(ImmutableBytesWritablekey, Result result, Context context) throwsIOException, InterruptedException
{
for(KeyValuekeyValue : result.list())
{
context.write(key, newKeyValue(key.get(), keyValue.getFamily(), keyValue.getQualifier(), HConstants.LATEST_TIMESTAMP,
KeyValue.Type.DeleteColumn));
}
}
}
After execution my hbase table looks like,

Code Walk Through:
Most of the code is self-explanatory, so you can easily check and get line by line understanding of the code.
I hope this program will help you to understand Bulk Deletion of column values using Hbase with hadoopmapreduce.
The purpose of sharing this post is to let you know how to perform bulk deletion of column values in hadoop development. It will be good if you practice and share your experience with proficient hadoop developers who have put their efforts and skills to intend this post.