Omar Gadir, Ph. D., Founder/CEO.
Source Of Unstructured Data
Unstructured data could be generated by machines (machine-made data) or by humans (man-made data). Machine data is generated by sensors, mouse clicks, and similar sources. It contains information about the activities and behavior of customers, transactions, servers, network, mobile devices, smart meters, log files, and weather. Data generated by humans is created on a daily basis and contains vital business information: interaction between employees, contracts, design, strategies, manuals, compliance documents, research, product development, marketing data, etc.
During the last few years, the focus of new products has been on machine data, which is broadly described as unstructured. Many articles, blogs, and discussions give the impression that machine data represents all unstructured data, which is not true. As a result, man-made data has been ignored, even though it is voluminous and contains valuable information. At Iteru, we believe that both machine-made and man-made data are important. To stress the importance of man-made data, I decided to write this blog.
Types of Data
There are two main types of data: structured and unstructured. Structured data refers to information with a high degree of organization, allowing seamless inclusion in a relational database. It is easy to extract information from structured data using data mining tools. Unstructured Data refers to information that does not have a pre-defined data model, resulting in irregularities and ambiguities that make it difficult to extract information using data mining tools. Experts estimate that 80 to 90 percent of the data in any organization is unstructured and is growing significantly, often faster than structured data.
There is also what is called semi-structured data. It does not conform to the formal structure of traditional data (like relational databases) but contains tags or other markers to separate semantic elements and enforce hierarchies within the data. It is also known as self-describing structure. Types of semi-structured data include: emails, XML, and JSON. For this document, semi-structured data is considered as part of unstructured data.
Machinemade data
Machine-made data is purely textual and has a simple predefined format. It often consists of delimited fields, such as those separated by commas (CSV) or spaces, making it easy to parse and extract columns. Machine data is frequently discarded after processing.
Man-made data
Man-made data has complex formats and is typically stored for a long time. It is more difficult to parse and process for several reasons:
- Irregular structure: Examples of documents with irregular structure include MS Word, PPT, PDF, audio, video, etc.
- Proprietary government and industry-based formats: Examples include the semiconductor industry’s Standard Test Data Format (STDF) and the pharmaceutical industry’s FASTQ format for biological sequences.
- Includes thousands of metadata that contain vital information: For instance, a PDF generates metadata that describes changes and functionality allowed within the document.
- Embedded documents: An MS Word document may include references to Excel documents or PDFs.
As an example of the complexity of man-made data, below is the format of a PDF file.
Current Analytics Solution
Most analytics products (e.g., Tableau, Platfora, Datameer) deal with three unstructured data types: flat files (CSV), Excel, and JSON. These three types constitute less than 1% of the total number of data types, while the majority are man-made.
To perform analytics, data must be transformed into flat files, Excel, and JSON using data extraction and mining algorithms. Many of these algorithms have several limitations:
- Data processed is often of poor quality, leading to inaccurate analytics.
- They cannot handle dynamic data; changes require re-execution of algorithms.
- They struggle to integrate conflicting or redundant data from diverse sources and formats.
- They are not efficient or scalable.
As a result, data mining algorithms have limited success when processing unstructured data, especially man-made data, leading to a large portion of unstructured data remaining unprocessed (dark data).
Unified Solution for Man-made and Machine-made Data
At Iteru, we believe both man-made and machine-made data are crucial. We have developed algorithms to address the challenges outlined above, including:
- Data cleansing.
- Dealing with dynamic data.
- Processing all 400 types of unstructured data.
- Providing efficient and scalable solutions.
Iteru's analytics and visualization tools cover all types of unstructured data, processing it and presenting results in a tabular form suitable for analytics tools (including third-party tools like Tableau, Platfora, and Datameer).