Model & evaluation
Measured results with the dataset and limits clearly stated.
Qwen2.5-VL or pretrained OCR extracts visual text. A TF-IDF logistic regression model suggests a category. An explicit skill rubric calculates coverage. Classifier accuracy is not extraction accuracy or hiring accuracy.
TF-IDF + logistic regression
The saved pipeline uses normalized skills with word unigrams and bigrams. The same feature transformation is used during training and inference. Training and test scenarios are specified in the CSV; no normalized skill sequence appears in both splits.
Dataset: 96 fictional skill-profile examples, authored for this mini project; 72 train and 24 held-out scenarios. No real candidate or employer data. Generalization is unproven.
Version: role-tfidf-v1 · scikit-learn 1.7.2
The evaluated model is saved after fitting only the training split. The test split is never included in fitting. These small fictional examples cannot establish general performance on real resumes.
| Role category | Precision | Recall | F1 | Test support |
|---|---|---|---|---|
| cloud-support | 1.0 | 1.0 | 1.0 | 3 |
| data-analyst | 0.75 | 1.0 | 0.857 | 3 |
| frontend | 1.0 | 0.667 | 0.8 | 3 |
| ml-intern | 1.0 | 1.0 | 1.0 | 3 |
| python-backend | 0.667 | 0.667 | 0.667 | 3 |
| qa-engineer | 1.0 | 0.667 | 0.8 | 3 |
| security | 1.0 | 1.0 | 1.0 | 3 |
| ui-ux | 0.75 | 1.0 | 0.857 | 3 |
Confusion matrix (rows = actual, columns = predicted)
Order: cloud-support, data-analyst, frontend, ml-intern, python-backend, qa-engineer, security, ui-ux
[3, 0, 0, 0, 0, 0, 0, 0] [0, 3, 0, 0, 0, 0, 0, 0] [0, 0, 2, 0, 0, 0, 0, 1] [0, 0, 0, 3, 0, 0, 0, 0] [0, 1, 0, 0, 2, 0, 0, 0] [0, 0, 0, 0, 1, 2, 0, 0] [0, 0, 0, 0, 0, 0, 3, 0] [0, 0, 0, 0, 0, 0, 0, 3]
Included sample extraction check
Engine: PDF text + RapidOCR. Five controlled fictional sample files; not an independent real-world benchmark. VLM has not been evaluated unless this report explicitly lists that engine.
| Sample | Skill precision | Skill recall | F1 |
|---|---|---|---|
| sample_resume.pdf | 1.0 | 1.0 | 1.0 |
| sample_resume.png | 1.0 | 0.75 | 0.857 |
| sample_phone_photo.jpg | 1.0 | 0.75 | 0.857 |
| sample_scanned_resume.pdf | 1.0 | 0.75 | 0.857 |
| sample_resume_v2.pdf | 1.0 | 1.0 | 1.0 |
These measure recognized skill names against manual annotations on the bundled fictional samples. They do not measure full text accuracy or certify VLM quality.
Reproduce the results
python train_model.py python evaluate_extraction.py --engine ocr
After configuring Ollama, use --engine vlm for a separate visual extraction evaluation. Results record the actual engine.