Andrew DeLisa's picture

Andrew DeLisa

ayan4m1

AI & ML interests

Distilled fine-tuning

Recent Activity

reacted to nwaughachukwuma's post with 🔥 about 5 hours ago
It’s easy to get distracted by benchmarks, throughput (tok/s), and all the hype around frontier model releases. This is Shiny Model Syndrome, which makes engineers and teams forget the basic physics of production software, i.e., using the right tool for the job and optimizing for ease of integration. - Teams spend huge amounts of money on frontier models for document parsing, OCR, detection, segmentation, and other task-specific visual AI workflows. - Inference marketplaces don’t find it profitable to list task-specific models like glm-ocr, paddleocr, or dots.mocr, even though they’re all superior to frontier VLMs for document parsing and OCR. - Engineers stitch together multiple endpoints for different use cases across the long tail of visual AI. Those who choose to self-host instead deal with painful infrastructure and GPU ops. At VLM Run, we wanted one place to run OCR models, VLMs, and ViTs that we could confidently use for our own internal agents and evals. The gateway was born out of that need, and we’ve since opened it to the public. The gateway exposes a single OpenAI-compatible endpoint for the long tail of visual AI across OCR, document parsing, VQA, detection, segmentation, embeddings, and transcription. Simply point the base_url of your OpenAI SDK at gateway.vlm.run/v1/openai, or ask your agent to connect via MCP (gateway.vlm.run/mcp). You can swap the model name to compare glm-ocr, dots.mocr, paddleocr-vl-1.6, qwen3.8-27b, gemma4-26b-a4b, and more. We handle serving, runtime, and pipelining behind the scenes to give you high-quality visual intelligence. - https://vlm.run/gateway - https://huggingface.co/blog/vlm-run/introducing-gateway - https://www.vlm.run/blog/introducing-gateway
View all activity

Organizations

All The Flavors's profile picture