Skip to content

Computer Vision & Multimodal AI Engineer

Read what a phone camera captures: a prescription, a scan, a diseased leaf held up in a field at midday.

Ref
J05
Team
Platform
Where
Remote, anywhere in India
All roles

A lot of HCF’s input arrives as a photograph taken on a cheap phone in bad light — a handwritten prescription, a leaf with early blight, a document nobody will retype. Making that legible to a machine is the difference between a tool that helps and a form nobody fills in.

You would work across NADI ai’s document and imaging pipeline and KisanMitra’s crop diagnosis.

What you would do

  • Build vision pipelines that tolerate glare, blur and bad framing
  • Work on document understanding for handwriting and non-standard layouts
  • Combine image and text signals where either alone is insufficient
  • Set honest accuracy thresholds, and refuse below them

What we are looking for

  • Python with PyTorch, OpenCV or similar
  • You have worked with images that were not a clean dataset
  • Understanding of OCR and document layout problems
  • Judgement about when a prediction is too weak to show a user

These are not calculator projects or website clones. You will work on RAG systems, real databases and production AI pipelines — the kind of work you can actually talk about in an interview.