Help and support
Creating Manual Models with Auto Data Extraction
While Artificio’s pre-trained AI models handle standard documents out of the box, you can easily xreate a custom model for unique, specialized document layouts. By combining Artificio’s Auto Data Extraction engine with manual labeling, you can precisely teach the AI exactly where to find your unique data fields and line-item tables.
Here is a step-by-step guide on how to build, label, and train a manual model.
Step 1: Set Up Your Manual Model
To get started, you need to create the container for your custom model and define the specific data points you want to extract.
Navigate to your Auto Data Extraction Tile and select Manual Model Creation option.
Add Model name and Define your Fields: * Global Labels: These are single-value fields that appear once per document (e.g., Invoice Number, Total Amount, Customer Name).
Table Labels: These are nested fields for line items that repeat across rows (e.g., Description, Quantity, Unit Price, Line Total).
Click Create Model.
Step 2: Use Auto Extraction & Manual Labeling
This is where the magic happens. Instead of drawing bounding boxes from scratch, Artificio uses its Auto Data Extraction engine to give you a massive head start.
1. Review Auto-Detected Fields
When you open an uploaded document in data view for the manually created model, Artificio automatically runs an initial OCR (Optical Character Recognition) scan. It will highlight text blocks it recognizes.
2. Labeling Global Fields
Click on a pre-defined field from your sidebar (e.g., Invoice Date).
Click or drag your mouse over the corresponding auto-extracted text on the document.
The system will map that specific location and text to your label.
3. Labeling Tables (Line Items)
Extracting tables manually can be tedious, but Artificio simplifies it into a few clicks:
Click the Add Table tool.
Define the boundaries of your table by dragging a box around the entire grid.
Map the Columns: Match your defined table labels (e.g., SKU, Price) to the columns in the document.
Verify Rows: The auto-extraction engine will attempt to separate the rows automatically. If a row is missed, simply click to add a row divider.
💡 Pro-Tip: If the auto-extraction misaligns a field due to poor document quality, you can manually redraw the bounding box or type the correct value directly into the validation panel.
Managing and Refining Your Model
Your manual model isn't static—it gets smarter over time.
Human-in-the-Loop Validation: As new documents pass through production, any low-confidence extractions will be flagged for your review.
Managing and Updating Your Models: To update a custom model, navigate to the Models list from your dashboard sidebar, locate your model, and click Edit. You can modify your configuration at any time by adding new labels, adjusting table structures, or uploading fresh training documents to continuously improve extraction accuracy.
Continuous Learning: When you correct a flagged document, that data can be fed back into the training pool, automatically updating the model to prevent the same error from happening twice.