Homepage Knowledge iOCR – Applications of Intelligent Character Recognition

Intelligent Automation and AI

iOCR – Applications of Intelligent Character Recognition

@mindbox

Zespół Mindbox

8 minutes

Processing visual stimuli is among the most complex and least understood processes in the human brain. It is known that people remember about 80% of what they see, and only 20% of what they read. Humans can process entire images seen by the eye in just 13 milliseconds[i] and are capable of remembering 2,000 images with at least 90% accuracy over several days, even with very short presentation times during learning[ii]. We can even see objects that do not exist[iii]. Algorithms are still very, very far from such efficiency and performance in image processing. Fortunately, from a business perspective, text is more important—it is the primary carrier of information, which is easier to extract from any image. This is handled by iOCR (Intelligent Optical Character Recognition) technology.    

What is iOCR software and what is it used for?

iOCR originates from OCR (Optical Character Recognition) technology developed in the early 1990s, which is a set of techniques and programs for recognizing characters and entire texts in raster graphic files. The task of OCR was usually to recognize individual letters of text in a scanned document[iv], most often printed in a specific typeface. After matching to the ASCII character table, a text file was created that represented the original content with varying degrees of accuracy. The main difficulty was, of course, reading handwriting, and some OCR programs even had trouble recognizing ligatures or characters printed in different fonts. Intelligent Character Recognition (iOCR) technology has gone much further—it manages to process images so that text files are prepared for machine processing by Robotic Process Automation (RPA) bots. Character conversion is based on intelligent pattern recognition. It is precisely the combination of iOCR with deep Machine Learning (ML) and Artificial Intelligence (AI) algorithms that has opened up entirely new perspectives for many organizations in terms of automating paper document processing. iOCR programs excel at handling documents with:
  • a fixed structure – e.g., tax forms, insurance forms, exam sheets, or documents;
  • a quasi-structure – such as invoices, purchase orders, payment confirmations, and shipping orders;
  • no structure – i.e., agreements, contracts, articles, letters, and other full-text documents.
Of course, the recognized document can also be a multi-page combination of these three types with plain text and graphic attachments. The importance of iOCR technology is confirmed by market research results[v], which indicate that as late as 2019, 49% of companies stored their business data in paper form, and 97% of the surveyed population was involved in digitizing their paper documents in one way or another. The conclusion is simple—using the latest achievements of iOCR is a necessary condition for carrying out digital transformation in an organization.

How does iOCR software work?

The process of recognizing the content of a bitmap and converting it into text can be divided into six stages:
  1. Scanning a paper document. The resulting bitmap is an electronic representation of the original. Of course, available utility software can control contrast, resolution, and colors, or invert colors, but the image remains understandable only to humans.
  2. Pre-processing. At this stage, the image is cleaned of noise, unnecessary areas outside the text are removed, and color and contrast are optimized to obtain an optimal bitmap for further processing. This is particularly important in handwriting processing. After pre-processing, an almost clean image of all characters is obtained, which improves text recognition results.
  3. Segmentation. The algorithm groups the processed image fragments into characters or character classes and searches for patterns matching the defined classes.
  4. Character extraction. In this key step, the search and recognition of characteristic features of individual letters take place. As a result, the previously extracted characters effectively become letters.
  5. Neural network training. Optionally, after extracting all the features of individual letters, characters can be fed into a neural network to teach it to recognize letters. The training dataset and methods used to achieve the best results will depend on the problem that requires an iOCR-based solution.
  6. Post-processing. This is where the results of the algorithms from the previous stages are refined and corrected. It is difficult to expect 100% effectiveness in text recognition because it depends to some extent on context, which is why the human eye makes the final corrections at the end.
As you can see, the key to intelligent, effective, and efficient image-to-text processing is primarily multi-layer artificial neural networks enabling Deep Machine Learning (DML). In practice, Recurrent Neural Networks (RNN – Recurrent Neural Networks)[vi], Long Short-Term Memory (LSTM – Long Short-Term Memory)[vii], and Convolutional Neural Networks (CNN – Convolutional Neural Networks)[viii] are most often used. Open-source libraries and tools such as OpenCV (Open Computer Vision Library)[ix], or the Vision API supported by Google, are also useful. The latter solution offers pre-trained models for extracting text from images of various types and quality.

What are the possible practical applications of iOCR technology?

iOCR works perfectly in any situation where information has been recorded in printed form and the organization needs that data in electronic form. The spectrum of applications is therefore practically unlimited. The first example that comes to mind is access control—after scanning an identity document, iOCR algorithms save the data in the appropriate database, and subsequent applications can check, for example, entry permissions. Filling out any paper forms becomes unnecessary. The same applies, for example, to car rental companies. Other areas where iOCR technology will certainly become a permanent fixture include finance departments, where, for example, purchase invoices are processed; research companies that will be able to automate survey processing; or the education sector, which will use iOCR to check exam sheets. The entire banking and insurance industry, where a great deal of paper documentation is processed, can use iOCR to automate, for example, credit application workflows, accident documentation, and insurance claim applications. One can also imagine iOCR operating in a bailiff’s office, where thousands of similar applications, summonses, and court letters are processed, as well as in sales departments for converting business cards into full-fledged records in a CRM system database.

What does an enterprise gain by implementing iOCR technology?

Thanks to iOCR, every enterprise gains two invaluable resources: money and time. If a company processes 50,000 purchase invoices per month[x], the iOCR program is able to read thousands of different invoice templates, not only Polish but also foreign ones, so the invoice entry time is very short. Additionally, if the program is confident in the data reading, it automatically generates an invoice file. In this way, about 20% of documents are created, and the rest go to a human, where iOCR has already filled in 80-85% of the fields. Every intervention by a human employee goes to machine learning algorithms, which improves the effectiveness of data recognition and reading over time. What are the consequences of this? No stress or fatigue for employees, especially at the end of the month when the largest portion of cost invoices arrives. This leads to fewer errors and faster invoice approval and payment processes. Fewer people are also needed for data entry, and personnel costs very often constitute the lion’s share of operating costs. It must be taken into account that humans are the main source of errors related to reading data and entering it into business systems. Therefore, having a digital and recognizable version of documents allows for the optimization of many business processes, accelerating their operation and improving reliability.

Why is it worth implementing iOCR technology in an enterprise?

Many mundane and repetitive processes in an organization can be automated and robotized thanks to intelligent character recognition—it is iOCR that prepares document versions so that they can be input material for software bots. Thanks to this, a company can much more effectively digitize, classify, store, and distribute documents of practically any type. This happens because iOCR combines the highest quality natural language processing, machine learning, and advanced character recognition functions to handle all types of business documents. The implementation of iOCR technology is a necessary condition for the robotization of business processes and the elimination of paper documents from the workflow. The speed of document processing indirectly improves overall work efficiency and reduces costs. In a situation where we are talking about tens of thousands of paper documents per month, financial savings related to reducing employment come to the fore. It is also worth implementing iOCR technology because it is easily integrated with other ERP or CRM business systems. Thanks to APIs, EDI, or databases, properly programmed bots can automate many other complex and advanced business processes, such as importing information from XML files into web applications or database records. Either way, iOCR is a key technology for implementing a business process digitization strategy. [i] https://pubmed.ncbi.nlm.nih.gov/24374558/ [ii] Cheryl L. Grady, Anthony R. McIntosh, M. Natasha Rajah and Fergus I. M. Craik, Neural correlates of the episodic encoding of pictures and words, PNAS 1998 March, 95 (5) 2703-2708. [iii] http://www.matematyka.wroc.pl/book/figury-niemo%C5%BCliwe [iv] https://pl.wikipedia.org/wiki/Optyczne_rozpoznawanie_znak%C3%B3w [v] https://info.aiim.org/capture-leaders-and-their-projects-2019?utm_campaign=20190424.webinar&utm_source=20190424w-Parascript1&utm_medium=20190424w-Parascript1 [vi] https://bulldogjob.pl/readme/3-typy-rekurencyjnych-sieci-neuronowych [vii] https://datascience.eu/pl/uczenie-maszynowe/zrozumienie-sieci-lstm/ [viii] http://www.cs.put.poznan.pl/alawrynowicz/SI_ML_CNN_2020_lawrynowicz.pdf [ix] https://github.com/opencv/opencv [x] https://impel.pl/blog/technologia-ocr-idealne-rozwiazanie-dla-faktur-kosztowych/

@mindbox

Zespół Mindbox

Newsletter

Subscribe to our Newsletter

Newsletter (EN)