Computer Vision & Multimodal AI Engineer
Read what a phone camera captures: a prescription, a scan, a diseased leaf held up in a field at midday.
- Ref
- J05
- Team
- Platform
- Where
- Remote, anywhere in India
All roles
A lot of HCF’s input arrives as a photograph taken on a cheap phone in bad light — a handwritten prescription, a leaf with early blight, a document nobody will retype. Making that legible to a machine is the difference between a tool that helps and a form nobody fills in.
You would work across NADI ai’s document and imaging pipeline and KisanMitra’s crop diagnosis.
What you would do
- Build vision pipelines that tolerate glare, blur and bad framing
- Work on document understanding for handwriting and non-standard layouts
- Combine image and text signals where either alone is insufficient
- Set honest accuracy thresholds, and refuse below them
What we are looking for
- Python with PyTorch, OpenCV or similar
- You have worked with images that were not a clean dataset
- Understanding of OCR and document layout problems
- Judgement about when a prediction is too weak to show a user
These are not calculator projects or website clones. You will work on RAG systems, real databases and production AI pipelines — the kind of work you can actually talk about in an interview.