AI with Papers - Artificial Intelligence & Deep Learning
описание
All the AI with papers. Every day fresh updates about #DeepLearning #MachineLearning #LLM & #ComputerVision Curated by Alessandro Ferrari | https://www.linkedin.com/in/visionarynet/ #AI #chatGPT
17 017
подписчиков
Охват к подписчикам
21,5%
ERR
Реакции к просмотрам
0,28%
299 на 24 постов
Пересылки к просмотрам
0,64%
678
Постов в день
0,0
всего 25
Где отзываются чаще
доля реакций к просмотрам- 10 июл.🔥ZipDepth: Depth on Any Device🔥 👉ZipDepth from UniBO is a super-compact monocular depth network by combining an efficient reparameterizable encoder-decoder with large-scale knowledge distillation from a foundation model. Repo under MIT💙 👉Review https://t.ly/qYrLZ 👉Paper https://arxiv.org/pdf/2607.08771 👉Project https://zipdepth.github.io/ 👉Repo https://github.com/fabiotosi92/ZipDepth0,70%
- 30 июл.🐠Dual-branch ID-Tracking🐠 👉TIDE: tracking dense, homogeneous targets, providing a scalable dual-branch design to accommodate diverse hardware constraints. MIT license💙 👉Review https://t.ly/WEDeY 👉Paper https://arxiv.org/pdf/2607.26412 👉Project https://vranlee.github.io/TIDE/ 👉Repo https://github.com/vranlee/TIDE0,54%
- 21 июл.👉Not a render. Not a concept. This is GENE.01 by Generative Bionics, the Italians coolest scaleup strikes back: in just six months, they turned GENE.01 into a fully functional humanoid platform that can walk, sense and interact. 👉Full-body multimodal skin perceives touch, proximity, force and temperature, bringing Physical AI closer to safe and natural collaboration with people. 👉More: https://t.ly/F3I3A0,54%
- 28 июл.🍿 Dawn of Generative Cinematography 🍿 🟩 The Odyssey by Christopher Nolan was shot entirely on IMAX 70mm. It feels almost romantic: massive cameras, film stock, premium lenses, and an obsessive pursuit of the highest possible quality at the moment of capture. 👉 Meanwhile, AI research is heading in the exact opposite direction. 🟩 A pre-print paper released today, "Camera Anything", demonstrates something that sounded like science fiction just a few years ago: you film a scene once... and then you can virtually reposition the camera anywhere. 👉More: https://t.ly/g-qUh 👉Paper arxiv.org/pdf/2607.24591 👉Project yixuanli98.github.io/cameraanything/ 👉Repo github.com/yixuanli98/CameraAnything0,46%
- 14 июл.🦧 MonkeyOCRv2 is out! 🦧 👉MonkeyOCRv2 is a text-centric visual foundation model that unifies fine-grained text modeling, cross-task representation learning, and cross-lingual generalization in a single encoder. Released for academic research and non-commercial use💙 👉Review https://t.ly/yicEK 👉Paper https://arxiv.org/pdf/2607.11562 👉Repo https://github.com/Yuliang-Liu/MonkeyOCRv20,43%
- 13 июл.🌔Foundation Global SFM🌔 👉Glob3R is a global SfM-style reconstruction built on 3D foundation models. key idea: explicitly optimize feed-forward geometric predictions. Repo TBA💙 👉Review https://t.ly/Z_4C7 👉Paper https://arxiv.org/pdf/2607.09225 👉Project https://junyuandeng.github.io/Glob3r/ 👉Repo TBA0,42%
- 7 июл.🐈⬛Spatial-perception native ViT🐈⬛ 👉LingBot-Vision, a vision foundation model pretrained to be spatial-perception native. Better than 7x bigger foundational models. Repo under Apache💙 👉Review https://t.ly/9xIso 👉Paper https://arxiv.org/pdf/2607.05247 👉Project https://technology.robbyant.com/lingbot-vision 👉Repo https://github.com/robbyant/lingbot-vision0,41%
- 24 июл.💢Unified Video Dense Prediction💢 👉UniD predicts: depth, surface normals, semantic segmentation, boundaries, human parts, albedo, shading, and materials. Code TBR💙 👉Review https://t.ly/oo7et 👉Paper https://arxiv.org/pdf/2607.21592 👉Project https://unid-video.github.io/ 👉Repo https://github.com/YihongSun/UniD0,40%
- 9 июл.🏵️SoccerNet 2026 Results🏵️ 👉The SoccerNet 2026 Challenges constitute the sixth annual edition of the SoccerNet open benchmarking effort, dedicated to advancing computer vision research in sports video understanding💙 👉Review https://t.ly/sfD4T 👉Paper https://lnkd.in/dSBgW_3s 👉Project https://lnkd.in/dfdmuvG80,38%
- 14 июл.🎂REMIND: long-term MOT re-ID🎂 👉REMIND by CVAR-UPM is a novel online tracker designed for long-term multi-object re-ID of generic indoor objects from monocular RGB, requiring neither camera pose nor depth. Repo under MIT💙 👉Review https://t.ly/AkQoI 👉Paper https://lnkd.in/dm58mkCv 👉Project https://lnkd.in/dZrAZqFe 👉Repo https://lnkd.in/dbidrwxU0,38%
- 29 июл.🔥Decoder-only Any-to-Any Model🔥 👉MODUS unifies any-to-any multimodal generation with one decoder, two experts, and zero task heads. Impressive work. Repo under Apache💙 👉Review https://t.ly/-2QKT 👉Paper https://lnkd.in/dhfBQGhB 👉Project https://lnkd.in/dPD_ECXk 👉Repo https://lnkd.in/dbDHw24u0,37%
- 11 июл.💋SAM-MT: Real-Time Multi-Target VOS💋 👉Fudan & Shangai unveil SAM-MT, an efficient interactive multi-target video segmentation framework that maintains near-single-object efficiency (FPS/VRAM) as target count increases, while maintaining robust video segmentation performance. Repo available💙 👉Review https://t.ly/Z_4C7 👉Paper https://lnkd.in/dvS-iyBD 👉Project https://lnkd.in/daQ8na8T 👉Repo https://lnkd.in/dgbX2tZv0,36%