Data Engineering - Real-time Data Streaming, Real-time Data Processing, Real-time Data Analytics | Data Engineering Solution in Bangalore | Apache Kafka Streaming Solutions in Bangalore | Kafka Confluent Cloud Solutions in Bangalore | Kafka Streaming Implementation Support in Bangalore | Apache Kafka Support in Bangalore | Multinode Kafka Cluster Setup in Bangalore | Kafka Application Consulting in Bangalore | Kafka cloud implementation in Bangalore | Kafka infrastructure consulting in Bangalore | Kafka security implementation in Bangalore | Kafka upgrade support in Bangalore | Zookeeper setup support in Bangalore | Zookeeper Solutions in Bangalore | Multinode Zookeeper Setup in Bangalore | Big Data Consulting Service Providers in Bangalore | Data Analytics Consutling Services in Bangalore | Big Data Solution Providers in Bangalore | Big Data Analytics Companies in Bangalore | Data Analytic Services in Bangalore | Big Data Services in Bangalore | Big Data Analytics Solutions in Bangalore | Big Data Analytics Service Providers in Bangalore | Big Data Case Studies | Big Data Companies in Bangalore | Multi Node Hadoop Cluster | Data Lake creation and support | Data Ingestion Services in Bangalore | Koolanch | Artificial Intelliegence Solutions in Bangalore | Predictive Analysis Solution in Bangalore | Machine Learning Solution in Bangalore | Deep Learning Solutions Bangalore | ChatBots for Websites | Text to Speech API | DialogFlow ChatBots | ChatBots using DialogFlow | AI based image processing | AI solution providers in Bangalore | AI based Predictive Analytics | Conversational Bots Development in Bangalore | AI chatbots and voicebots | E-Commerce Solution Providers in Bangalore | Demandware Consulting Service in Bangalore | Demandware Companies in Bangalore | SFCC Consulting Service in Bangalore | SFCC Consulting Companies in Bangalore | SFCC Service Providers in Bangalore | Demandware Contract Staffing in Bangalore | Salesforce Commerce Cloud Consulting Services in Bangalore | SFCC Contract Staffing in Bangalore | Salesforce Commerce Cloud Contract Staffing in Bangalore | Oracle Consulting Services in Bangalore | Oracle Service Providers in Bangalore | Oracle Contract Staffing in Bangalore | OCC Contract Staffing in Bangalore | Oracle Commerce Cloud Consulting in Bangalore | Oracle Commerce Cloud Companies in Bangalore | SAP Hybris Consulting Services in Bangalore | SAP Hybris Service Providers in Bangalore | SAP Hybris Contract Staffing in Bangalore | SAP Hybris Commerce Cloud Consulting in Bangalore | SAP Hybris Companies in Bangalore | SAP Hybris Solutions in Bangalore | Hybris Commerce Solution in India | Hybris Solution Provider Companies | Magento Consulting Services in Bangalore | Magento Service Providers in Bangalore | Magento Contract Staffing in Bangalore | Magento Commerce Cloud Consulting in Bangalore | Magento Companies in Bangalore | Mobile App Development Company in Bangalore | Android App Development Services in Bangalore | Location Tracking Based Mobile App Development | Mobile App Development In Bangalore | Mobility Solution Provider in Bangalore | SQL Server Support Services in Bangalore | SQL Server Support Companies in Bangalore | Data Mining Solution in Bangalore | Custom App Development in Bangalore | Contract Staffing Solution in Bangalore

02JunJune 2, 2018

Processing and Analysis of Big Telecom Data to minimize crime, combat terrorism, unsocial activities etc.

Telecom providers have a treasure trove of captive data - customer data, CDR (call detail records), call center interactions, tower logs etc. and are metaphorically “sitting on a gold mine”. Ideally, each category of the generated data has the following information. ⦁ Customer data consolidates customer id, plan details, demographic, subscribed services and spending patterns ⦁ Service data category consolidates types of customer, customer history, complain category, query resolved etc. are on ⦁ Usually for the smart mobile phone subscriber,...

By Kislay KomalApache Hadoop, Architecture, Data Analysis, Data EngineeringBig Data Processing and Analysis, Processing and Analysis of Big Telecom Data, Processing and Analysis of Telecom Data, Telecom Data, Telecom Data Processing and AnalysisComments Off

21JanJanuary 21, 2018

Deleting Solr log files/folder from Standby NameNode could be the disaster when Primary NameNode is active in the HDP (Hortonworks Data Platform) Hadoop Cluster

Most of us know that we use Apache Ambari for managing, provisioning and monitor different components of a Hortonworks Hadoop cluster. We also know that Apache Ranger can be used as a centralized security administration solution for Hadoop that enables administrators to create and enforce security policies for HDFS and other Hadoop platform components. When ranger hdfs plugin is enabled ,it writes the client interaction activity to Solr if it is configured. The default location of this solr log files...

By Kislay KomalApache Hadoop, Data EngineeringComments Off

12DecDecember 12, 2017

Fault Tolerance Enhancement On Apache Hadoop 3.0.0-alpha2 For Supporting More Than 2 NameNodes

NameNode is the most critical resource in Hadoop core cluster. Once very large files loaded into the Hadoop Distributed File System (HDFS), the files get broken into block-sized chunks as per the parameter configured (64 MB by default). The chunks are then stored as independent units across the data nodes in the cluster. The primary responsibility of the data nodes is to hold the actual data in the form of chunk and NameNode holds the information where all the chunks located/stored in the...

By Kislay KomalApache Hadoopactive NameNode, Apache Hadoop, Apache Hadoop 3.0.0, Apache Hadoop 3.0.0-alpha2, configuration of NameNodes, data nodes, data nodes location, Fault Tolerance Enhancement On Apache Hadoop, filesystem metadata, filesystem namespace, filesystem tree, FsImage, FsImage file, Hadoop 2.0.0, Hadoop 3.0.0, Hadoop core cluster, Hadoop Distributed File System, Hadoop framework, HDFS, HDFS cluster, HDFS NameNode, JournalNodes, Master Node, NameNode, NameNode holds the information, new features of Apache Hadoop 3.0.0, new features of Hadoop 3.0.0, Primary NameNode, Quorum-based Storage, responsibility of the data nodes, secondary NameNode, single point of failure, SPOF, standby NameNode, Supporting More Than 2 NameNodes, very large amount of enterprise dataComments Off

30NovNovember 30, 2017

Basic Understanding Of Stateful Data Streaming Supported By Apache Flink

Technologies related to Big Data processing platform are enhancing the maturity in order to efficiently execute the streaming data which is becoming a major focus point to take business decision instantly specially in telecom and retail sector. Collecting data continuously from the various sensors installed/fitted with an industrial heavy equipment, click stream on an e-commerce application’s navigation etc can be considered as streaming data generation sources. By leveraging streaming application, we can process/analyze these continues flow of data without...

By Kislay KomalData Engineering, Processing EngineApache Flink, Big Data processing platform, checkpoint, concurrent updates, Data Streaming, Flink, HDFS, KeyBy, maintain parallelism, protecting streaming application from failure, savepoint., Stateful Data Streaming, Stateful Data Streaming Supported By Apache Flink, stateful operation, Stateless computation, Statelful computation, streaming application, Understanding Of Stateful Data Streaming Supported By Apache FlinkComments Off

25SepSeptember 25, 2017

Apache Flink – A 4G Data Processing Engine

Analyzing streaming data in large-scale systems is becoming a focal point day by day to take accurate business decisions due to mushrooming of digital data generation sources around the globe including social media. Real-Time analytics are becoming more attractive due to possibilities of getting insights from the time-value of data (in other words, when data is in motion). Apache Flink, an open source highly innovative stream processor engine has been grounded which helps to take advantage of stream-based approaches. Besides...

By Gautam GoswamiData Engineering, Processing Engine4G Data, 4G Data Processing, 4G Data Processing Engine, Amazon EC2, analyze historical data, Apache, Apache Flink, Big Data, core computational belt, data pipeline, Data Processing Engine, Data stream source, dataArtisans, DataSet API, DataStream API, Delta Iterate, DGI, Directed Acyclic Graph, disk spilling, fault-tolerant, flatMap, Flink, Flink Runtime, Garbage Collector, Google cloud, GroupReduc, Hadoop, hashing and sorting, Iterate, Java Virtual Machine, JVM, JVM memory management system, map, memory management system, micro-batching, multi-node clusters, processor engine, real-time analytics, Savepoints, single JVM, Single node clusters, storage mechanism, stream processor, stream processor engine, Twitter streaming, Yahoo, YARN, Yet Another Resource NegotiatorComments Off

17SepSeptember 17, 2017

Steering number of mapper (MapReduce) in sqoop for parallelism of data ingestion into Hadoop Distributed File System (HDFS)

To import data from most the data source like RDBMS, sqoop internally use mapper. Before delegating the responsibility to the mapper, sqoop performs few initial operations in a sequence once we execute the command on a terminal in any node in the Hadoop cluster. Ideally, in production environment, sqoop installed in the separate node and updated .bashrc file to append sqoop's binary and configuration which helps to execute sqoop command from anywhere in the multi-node cluster. Most of the...

By Gautam GoswamiData Engineering, Data IngestionData ingestion, Hadoop Distributed File System, HDFS, Map Reduce, parallelism of data ingestion, Sqoop, sqoop for parallelism of data ingestion into Hadoop Distributed File System (HDFS)Comments Off

08SepSeptember 8, 2017

Transfer structured data from Oracle to Hadoop storage system

Using Apache's sqoop, we can transfer structured data from Relational Database Management System to Hadoop distributed file system (HDFS). Because of distributed storage mechanism in Hadoop Distributed File System (HDFS), we can store any format of data in huge volume in terms of capacity. In RDBMS, data persists in the row and column format (Known as Structured Data). In order to process the huge volume of enterprise data, we can leverage HDFS as a basic data lake. In this...

By Gautam GoswamiData Engineering, Hadoop Eco SystemAmazon web service, Apache Sqoop, Apache's sqoop, Data ingestion, Data ingestion mechanism, distributed storage, distributed storage mechanism, enterprise data, Google cloud, Hadoop, Hadoop 2.x, Hadoop Distributed File System, Hadoop storage system, HDFS, huge volume of enterprise data, Microsoft Azure, multi node cluster, Oracle to Hadoop, Sqoop, structured data, Transfer structured data from Oracle to Hadoop, Using Apache sqoopComments Off

29AugAugust 29, 2017

Data Ingestion phase for migrating enterprise data into Hadoop Data Lake

The Big Data solutions helps to achieve valuable information to iron out the accurate strategic business decision. Exponential growth of digitalization, social media, telecommunication etc. are fueling enormous data generation everywhere. Prior to process of huge volume of data, we should have efficient data storage mechanism in a distributed manner to hold any form of data starting from structured to unstructured. Hadoop distributed file systems (HDFS) can be leveraged efficiently as data lake by installing on multi node cluster....

By Gautam GoswamiData Engineering, Data IngestionApache software foundation, Apache Sqoop, ata storage mechanism, ATG database, ATG database schema, cloud service providers, collecting Twitter streaming data, Couchbase, Data ingestion, Data Ingestion phase for migrating enterprise data into Hadoop Data Lake, Data Lake, data storage mechanism, DB2, Digitization, distributed storage, efficient data storage mechanism, ELT, enterprise data, export data from Kafka topic to HDFS, fault-tolerant, Flume, Hadoop, HADOOP Cluster, Hadoop Data Lake, Hadoop distributed file systems, Hadoop multi node cluster, HDFS, Hive, huge data reservoirs, huge volume of data, Ingestion, JDBC connector, JDBC protocol, Kafka, Kafka HDFS connector, Kafka to HDFS, Mainframe, mainframe dataset to HDFS, MapReduce, MapReduce distributed computing, migrating enterprise data, moving large amount of streaming data into HDFS, multi node cluster, multiple delimited text files, MySQL, Netezza, NoSql DB, NoSql Stores, Oracle, Oracle 11g Enterprise Edition, Oracle ATG Platform, parallel import process, parallel processing, pluggable mechanism, PostgreSQL, read the messages from Kafka topic, SQLServer, Sqoop, Sqoop installation, Strom, Using Kafka HDFS connectorComments Off

03JulJuly 3, 2017

Big Data Analytics in Banking Systems

Typically Banking systems are responsible to validate and verify financial transaction data, geo-location data from mobile devices, merchant data, and authorization including submission data. Data from lots of social media channels and Banking’s mainframe data center have a significant challenge to process and deliver final output. The Issue: Legacy systems are incapable of processing the data in when is in motion. Combining all different format of data is together is another challenge like structured, semi- structured and un-structured. Big data Approach:- Big data analytics...

By Kislay KomalData EngineeringBig Data, Big Data Analysis, Big Data Analytics, Big Data Analytics and Banking Systems, Big Data Analytics in Banking, Big Data Analytics in Banking Systems, Big Data in Banking SystemsComments Off

08JunJune 8, 2017

Why Lambda Architecture in Big Data Processing

Due to the exponential growth of digitization, the entire globe is creating minimum 2.5 Quintilian 2500000000000 Million) bytes of data every day and that we can denote as Big Data. Data generation is happening from everywhere starting from social media sites, various sensors, satellite, purchase transaction, Mobile, GPS signals and much more. With the advancement of technology, there is no sign of slowing down of data generation, instead it will grow in massive volume. All the major organizations, retailers,...

By Gautam GoswamiArchitecture, Data EngineeringApache Kafka, Apache Spark, Batch Data-processing Pipeline, batch layer, Big Data, Big Data Processing, big data technologies, Data ingestion, data pipeline, Data Pipelines, data processing pipeline, data warehousing systems, design framework, Digitization, fault tolerance, Flume Lambda sign λ, GPS signals, Hadoop Distributed File System, HDFS, Lambda, Lambda Architecture, Lambda Architecture in Big Data Processing, Lambda Architecture is a pluggable architecture, leveraging big data technologies, live streaming data, Mobile, Nathan Marz, persistence of data, pluggable architecture, purchase transaction, Quintilian bytes of data, satellite, sensors, Serving Layer, Speed layer, Streaming Data Pipeline, streaming data processing pipeline, streaming layer, Streaming or Speed layerComments Off