Skip to main content
Browse help topics

Help and support

Introduction

Imagine this: Every day, you spend time copying and pasting data from various sources like websites, emails, or spreadsheets. It's tedious and prone to errors. 

Data extraction automation is like having a tireless assistant who takes over that repetitive task. It sets up an automatic process to: 

  • Connect to your data sources: This could be a website, database, email, or even another spreadsheet. 

  • Find the specific data you need: The automation can identify the exact information you're looking for, like product prices on a website or customer contact details in emails. 

  • Extract the data: Just like magic, the data is automatically pulled out and collected. 

  • Deliver the data: This could be placing it in a report, spreadsheet, or database, wherever you need it most. 

Benefits of data extraction automation

  • Save you time: Focus on more strategic tasks while the automation handles the data collection. 

  • Reduce errors: No more typos or missed information from manual copying. 

  • Get data faster: Access real-time information for better decision-making. 

  • Improve accuracy: Consistent and reliable data leads to better analysis. 

Think of it as putting your data collection on autopilot!  

Automate Your Document Processing! 

To get started with automating your document processing, we'll create a dataflow. This dataflow acts like a recipe that tells our system how to handle your documents. 

Building Your Dataflow

  1. Give it a Name: Choose a descriptive name for your dataflow, up to 45 characters. Avoid names starting with "auto_" as we reserve those for internal use. 

  1. Describe what it Does: Briefly explain what kind of documents this dataflow will process and what information it will extract. 

  1. Select Email Source (Optional): If your documents arrive via email, choose the email address from where you want to fetch them. 

Data Processing Options

  • Fixing Skewed Documents: Sometimes documents get scanned at an angle. This option straightens them for easier processing. 

  • Setting Orientation: Indicate whether documents are typically portrait or landscape to ensure proper reading. 

  1. Data Group Selection: Choose the group of data elements (like customer information, invoice details) you want to extract from the document. 

  1. Application Flow: This is where the magic happens! Select the processing steps you need based on the document type and desired outcome. Here's a breakdown of the options: 

  • 20-AI OCR (Optical Character Recognition): This extracts text from the document like names, addresses, or amounts. 

  • 21-AI Sense: This intelligent feature understands the overall context of the document and identifies relevant data points for extraction. 

  • 22- Signature Detection: This identifies and locates any signatures present in the document. 

  • 30- Classification: This categorizes the document type (e.g., invoice, receipt, contract) for further processing. 

  • 33- Selection Group: This extracts data from predefined areas within the document you've identified as relevant. 

  • 35- Table Classification: This identifies and organizes data presented in tabular format within the document. 

  • 40- Entity Extraction: This extracts specific pieces of information from the document, like names, dates, or product codes. 

  • 70- Data Verification: This double-checks the extracted data for accuracy before proceeding. 

  • 80- Data Validation: This confirms that the extracted data meets your defined criteria (e.g., date format, value range). 

  • 95- Email Transmission: If needed, this step sends an email based on the processed document data. 

  • 100- Integration Model: This allows you to integrate the extracted data with other systems you use, streamlining your workflows. 

By creating a dataflow with these options, you can automate document processing, saving your time and ensuring accurate information extraction!  

After uploading your documents, you can click on the "Dataflow Result" tab to track the progress of your analysis. This tab provides a visual representation of what's happening to your documents behind the scenes. 

Here's what you'll see: 

  • Processing Stages: You'll see a series of steps representing the different analyses your document is undergoing. These steps may vary depending on the dataflow you selected and the models you trained. Common examples include:  

  • NER Prediction (Named Entity Recognition): This stage identifies and classifies important entities within your document, such as names, locations, or dates. 

  • Table Classification: This stage identifies and categorizes any tables found in your document. 

  • Signature Classification: This stage identifies and classifies any signatures present in your document. 

  • Progress Indicators: Each stage will have a progress indicator that shows you how far along it is in the analysis.  

Essentially, the dataflow result tab gives you a clear picture of how your documents are being analyzed and helps you understand what's happening at each stage of the process. 

Additional Points: 

  • The specific stages and their order may differ depending on the dataflow you selected. 

  • Once all processing stages are complete, the results will be displayed within the data view interface. 

I hope this explanation clarifies the functionality of the dataflow result tab! 

Conclusion

Data extraction automation saves time and effort by automatically gathering data. It improves accuracy and efficiency by eliminating manual errors and providing real-time information. You can customize data extraction through dataflows to get the information you need.