apple/LensVLM-9B · Hugging Face
LensVLM is a 9B Vision Language Model (VLM) that scans compressed images of text,
then selectively expands only the relevant pages to their uncompressed form via
learned tools.
Paper: LensVLM: Selective Context Expansion for Compressed Visual Representation of Text
Code: https://github.com/apple-aiml-research/ml-lensvlm
License
All ML model files in this repository, including Apple's modifications to the Qwen
model, are provided under the terms of the
Apple Machine Learning Research Model Licens...
Read more at huggingface.co