ACADEMIC DEMONSTRATION

Model & evaluation

Measured results with the dataset and limits clearly stated.

Three separate components

Qwen2.5-VL or pretrained OCR extracts visual text. A TF-IDF logistic regression model suggests a category. An explicit skill rubric calculates coverage. Classifier accuracy is not extraction accuracy or hiring accuracy.

Held-out classifier accuracy87.5%
Macro F10.873
Training examples72
Held-out examples24

TF-IDF + logistic regression

The saved pipeline uses normalized skills with word unigrams and bigrams. The same feature transformation is used during training and inference. Training and test scenarios are specified in the CSV; no normalized skill sequence appears in both splits.

Dataset: 96 fictional skill-profile examples, authored for this mini project; 72 train and 24 held-out scenarios. No real candidate or employer data. Generalization is unproven.

Version: role-tfidf-v1 · scikit-learn 1.7.2

The evaluated model is saved after fitting only the training split. The test split is never included in fitting. These small fictional examples cannot establish general performance on real resumes.

Role categoryPrecisionRecallF1Test support
cloud-support1.01.01.03
data-analyst0.751.00.8573
frontend1.00.6670.83
ml-intern1.01.01.03
python-backend0.6670.6670.6673
qa-engineer1.00.6670.83
security1.01.01.03
ui-ux0.751.00.8573
Confusion matrix (rows = actual, columns = predicted)

Order: cloud-support, data-analyst, frontend, ml-intern, python-backend, qa-engineer, security, ui-ux

[3, 0, 0, 0, 0, 0, 0, 0]
[0, 3, 0, 0, 0, 0, 0, 0]
[0, 0, 2, 0, 0, 0, 0, 1]
[0, 0, 0, 3, 0, 0, 0, 0]
[0, 1, 0, 0, 2, 0, 0, 0]
[0, 0, 0, 0, 1, 2, 0, 0]
[0, 0, 0, 0, 0, 0, 3, 0]
[0, 0, 0, 0, 0, 0, 0, 3]

Included sample extraction check

Engine: PDF text + RapidOCR. Five controlled fictional sample files; not an independent real-world benchmark. VLM has not been evaluated unless this report explicitly lists that engine.

SampleSkill precisionSkill recallF1
sample_resume.pdf1.01.01.0
sample_resume.png1.00.750.857
sample_phone_photo.jpg1.00.750.857
sample_scanned_resume.pdf1.00.750.857
sample_resume_v2.pdf1.01.01.0

These measure recognized skill names against manual annotations on the bundled fictional samples. They do not measure full text accuracy or certify VLM quality.

Reproduce the results

python train_model.py
python evaluate_extraction.py --engine ocr

After configuring Ollama, use --engine vlm for a separate visual extraction evaluation. Results record the actual engine.