Bulk CV upload and parsing
Bulk CV upload and parsing that lands in real records.
A folder of CVs becomes a queue, each one is read into text and a proposed candidate, and a recruiter approves every record before it exists.
What changes
-
Forty CVs in, forty records out, checked
Name, title, email, phone, location and years of experience are read off the top of each CV and shown beside the line they came from.
-
No duplicates from a bulk upload
The proposed email is checked against your candidates before anything is created. A match shows the existing person, and filing under them only fills blanks.
-
Findable from the first minute
The CV text is on the record and in search from the moment it is approved, and the skills in your own vocabulary are suggested as tags.
Start to finish
From a folder to the pool.
-
Choose the files
PDF, Word or plain text, up to 40 in a batch and 15MB each. Anything else is refused by name before it is uploaded, and the reason is shown.
-
They go straight to storage
Each file gets its own signed upload, and its place in the queue is created first. A file that uploads and then loses its browser tab is still in the queue, not lost in a bucket.
-
Each one is read into text
Word documents from their own structure, PDFs from their text layer. The result carries a confidence: high, good, low, or none with the reason, such as a scan.
-
A candidate is proposed
The header block of a CV is a name, a title, a phone, an email and a suburb in some order. Those are read out, with a LinkedIn address and years of experience where the CV states them in words, each beside its source line.
-
Duplicates are caught first
Before creation, the email is matched against your existing candidates. A match shows the person you already hold. Filing the CV under them fills empty fields only.
-
Tags are suggested, not assumed
The text is searched for the terms in your controlled vocabulary, synonyms included. Each suggestion is a tick box with the sentence it came from, and it is stored as read from a document until a recruiter verifies it.
-
You approve, or discard
Approving creates the candidate with the source recorded, files the CV as their primary document and puts the text into search. Discarding removes the file from the queue and from storage.
Formats and limits
What is read, what is refused, and what it tells you.
Every extraction carries a confidence and a reason. A half-read CV that looks whole is worse than a refusal, because the missing half is invisible forever.
| File | What happens |
|---|---|
| Word (.docx) | Read from the document’s own structure. Paragraphs, line breaks and tab columns are kept, so a two-column skills table stays readable. |
| PDF with a text layer | Read from the text content. This covers PDFs produced by Word, Google Docs, most applicant tracking systems and LaTeX. |
| Plain text (.txt, .md) | Read as is. |
| Scanned or photographed PDF | Recognised as a scan and reported as one. No text is invented. The screen asks for the text to be pasted or for a Word version. |
| PDF whose fonts carry no usable encoding | The text comes out unreadable. The check that catches a scan catches this too, and the CV is reported rather than stored as gibberish. |
| Old Word (.doc) | Refused with a message asking for a .docx or a PDF. |
-
40 files per batch
A real folder of CVs. A file server’s worth belongs in an import, not four hundred uploads.
-
15MB per file
Larger than any CV, small enough that a mistaken video never reaches the reader.
-
200,000 characters kept
The text stored on the record. A CV runs to a few thousand; the cap only matters for a 400-page PDF, and it truncates rather than refuses.
-
Six read at a time
Extraction runs in the background in pages, so the review of each CV opens instantly and one unreadable file never stalls the rest.
Controls
The rules the reader follows.
The machine proposes and a person signs. That is the same rule as AI summaries, CV redaction and opportunity requirements, and it matters most here.
-
Nothing created on its own
Every candidate from an upload was approved by a recruiter on a screen that showed the evidence.
-
No guess where nothing was found
An empty field is a recruiter typing one thing. A confidently wrong field is a recruiter reading past it.
-
Nothing overwritten
A CV is evidence about a person, not a replacement for what a recruiter learned on the phone. Filing under an existing record fills blanks only.
-
Years of experience only when stated
Never computed from the earliest date on the page. A graduate with a part-time job from 2014 is not a candidate with twelve years’ experience.
Related
Questions
Which CV file formats can be parsed?
PDF, Word in .docx format, and plain text. A PDF is read from its text layer, so a PDF produced by Word, Google Docs or another system reads well. An old .doc file is refused with a message asking for a .docx or a PDF.
What happens with a scanned CV?
It is recognised as a scan or a photograph and reported as one, rather than returned as an empty record that looks read. There is no optical character recognition. The screen asks for the text to be pasted or for a Word version.
How accurate is the parser?
We do not publish an accuracy figure, because we have not measured one we would stand behind. What the product does instead is show every proposed field beside the line of the CV it was read from, so a wrong value is visible before it is saved. Where nothing is recognised, the field is left empty rather than guessed.
How many CVs can I upload at once?
Up to 40 files in a batch, each up to 15MB. Files go straight to storage, one signed upload each, so a large folder does not fail part-way through. Reading runs six at a time in the background until the queue is done.
Does it create candidate records automatically?
No. Extraction proposes a record; a recruiter approves it. A wrong email address on a candidate record is not a typo, it is a stranger receiving somebody else’s job applications, so nothing is created without a person saying yes.
What if the person is already in our database?
Before anything is created, the proposed email address is checked against your existing candidates. If there is a match you see the existing record and can file the CV under it. Filing under an existing person only fills blank fields; it never overwrites what a recruiter has already recorded.
Is the text of the CV searchable afterwards?
Yes. The extracted text is stored on the document and included in the candidate’s search text from the moment the record is created, so a search for a skill finds the person before anybody has tagged them.
Does a language model read the CV?
No. The field reader and the tag matcher are deterministic rules: a set of patterns for names, titles, contact details and locations, and a search for the terms in your own controlled vocabulary. The same rules run on every document, and the result can be checked line by line.
Ready when you are
See Recruited on one of your live roles.
A demo against a real role tells you more than a feature list. Bring one, and we will walk it through the platform with you.
We reply within one Australian business day.