tgindex

AI with Papers - Artificial Intelligence & Deep Learning

описание

All the AI with papers. Every day fresh updates about #DeepLearning #MachineLearning #LLM & #ComputerVision Curated by Alessandro Ferrari | https://www.linkedin.com/in/visionarynet/ #AI #chatGPT

17 014
подписчиков

Лучшие посты

за три месяца
  • 7 июл.4 119 просмотров17 реакций35 пересылок

    🐈‍⬛Spatial-perception native ViT🐈‍⬛ 👉LingBot-Vision, a vision foundation model pretrained to be spatial-perception native. Better than 7x bigger foundational models. Repo under Apache💙 👉Review https://t.ly/9xIso 👉Paper https://arxiv.org/pdf/2607.05247 👉Project https://technology.robbyant.com/lingbot-vision 👉Repo https://github.com/robbyant/lingbot-vision

  • 10 июл.4 026 просмотров28 реакций49 пересылок

    🔥ZipDepth: Depth on Any Device🔥 👉ZipDepth from UniBO is a super-compact monocular depth network by combining an efficient reparameterizable encoder-decoder with large-scale knowledge distillation from a foundation model. Repo under MIT💙 👉Review https://t.ly/qYrLZ 👉Paper https://arxiv.org/pdf/2607.08771 👉Project https://zipdepth.github.io/ 👉Repo https://github.com/fabiotosi92/ZipDepth

  • 15 июл.4 006 просмотров12 реакций31 пересылок

    🌈FlowWAM: flow->action prediction🌈 👉FlowWAM is a novel dual-stream diffusion framework that adopts optical flow as a unified, video-native action representation. Repo under Apache💙 👉Review https://t.ly/FmutT 👉Paper https://arxiv.org/abs/2607.13017 👉Project https://flow-wam.github.io/ 👉Repo github.com/YixiangChen515/FlowWAM

  • 14 июл.3 959 просмотров17 реакций38 пересылок

    🦧 MonkeyOCRv2 is out! 🦧 👉MonkeyOCRv2 is a text-centric visual foundation model that unifies fine-grained text modeling, cross-task representation learning, and cross-lingual generalization in a single encoder. Released for academic research and non-commercial use💙 👉Review https://t.ly/yicEK 👉Paper https://arxiv.org/pdf/2607.11562 👉Repo https://github.com/Yuliang-Liu/MonkeyOCRv2

  • 11 июл.3 849 просмотров14 реакций35 пересылок

    💋SAM-MT: Real-Time Multi-Target VOS💋 👉Fudan & Shangai unveil SAM-MT, an efficient interactive multi-target video segmentation framework that maintains near-single-object efficiency (FPS/VRAM) as target count increases, while maintaining robust video segmentation performance. Repo available💙 👉Review https://t.ly/Z_4C7 👉Paper https://lnkd.in/dvS-iyBD 👉Project https://lnkd.in/daQ8na8T 👉Repo https://lnkd.in/dgbX2tZv

  • 17 июл.3 736 просмотров7 реакций16 пересылок

    🏯SOTA Music-to-Dance Gen🏯 👉The Tongyi Lab unveils Wan-Dancer, a novel stable minute-scale synthesis at 720p/30fps across five dance genres. Impressive results, new SOTA on long clip by a large margin. Repo under Apache 2.0💙 👉Review https://t.ly/AKY5j 👉Paper https://lnkd.in/d_xA7dwb 👉Project https://lnkd.in/dzfnw2h4 👉Repo https://lnkd.in/d-Zj_cTf

  • 14 июл.3 664 просмотров14 реакций39 пересылок

    🎂REMIND: long-term MOT re-ID🎂 👉REMIND by CVAR-UPM is a novel online tracker designed for long-term multi-object re-ID of generic indoor objects from monocular RGB, requiring neither camera pose nor depth. Repo under MIT💙 👉Review https://t.ly/AkQoI 👉Paper https://lnkd.in/dm58mkCv 👉Project https://lnkd.in/dZrAZqFe 👉Repo https://lnkd.in/dbidrwxU

  • 9 июл.3 651 просмотров14 реакций29 пересылок

    🏵️SoccerNet 2026 Results🏵️ 👉The SoccerNet 2026 Challenges constitute the sixth annual edition of the SoccerNet open benchmarking effort, dedicated to advancing computer vision research in sports video understanding💙 👉Review https://t.ly/sfD4T 👉Paper https://lnkd.in/dSBgW_3s 👉Project https://lnkd.in/dfdmuvG8

  • 13 июл.3 596 просмотров15 реакций27 пересылок

    🌔Foundation Global SFM🌔 👉Glob3R is a global SfM-style reconstruction built on 3D foundation models. key idea: explicitly optimize feed-forward geometric predictions. Repo TBA💙 👉Review https://t.ly/Z_4C7 👉Paper https://arxiv.org/pdf/2607.09225 👉Project https://junyuandeng.github.io/Glob3r/ 👉Repo TBA

  • 23 июл.3 368 просмотров9 реакций21 пересылок

    🫛Spatially-Aware Class-Agnostic Counting🫛 👉UpCount is reference-free, spatially aware, class-agnostic object counting with an MAE-pretrained ViT, DPT-style feature reassembly, FeatUp-style joint bilateral upsampling & proposal verification. Repo MIT💙 👉Review https://t.ly/dWOc3 👉Paper https://arxiv.org/pdf/2607.16826 👉Repo github.com/r28112072-rgb/upcount

  • 24 июл.3 271 просмотров13 реакций27 пересылок

    💢Unified Video Dense Prediction💢 👉UniD predicts: depth, surface normals, semantic segmentation, boundaries, human parts, albedo, shading, and materials. Code TBR💙 👉Review https://t.ly/oo7et 👉Paper https://arxiv.org/pdf/2607.21592 👉Project https://unid-video.github.io/ 👉Repo https://github.com/YihongSun/UniD

  • 22 июл.3 214 просмотров6 реакций14 пересылок

    🦜Streaming 4D Transformer🦜 👉IGGT4D is a novel a streaming instance-grounded geometry transformer for online 4D scene understanding. It processes video frames sequentially, reuses historical context through causal spatial-temporal modeling, and incrementally updates a unified representation of camera motion, geometry, and object identity. Repo/Data announced💙 👉Review https://t.ly/LFrKR 👉Paper https://arxiv.org/pdf/2607.19228 👉Project https://iggt4d.github.io/ 👉Repo TBA

  • 21 июл.3 024 просмотров3 реакций

    What about more posts about Robotics?

  • 21 июл.2 963 просмотров16 реакций18 пересылок

    👉Not a render. Not a concept. This is GENE.01 by Generative Bionics, the Italians coolest scaleup strikes back: in just six months, they turned GENE.01 into a fully functional humanoid platform that can walk, sense and interact. 👉Full-body multimodal skin perceives touch, proximity, force and temperature, bringing Physical AI closer to safe and natural collaboration with people. 👉More: https://t.ly/F3I3A

  • 28 июл.2 884 просмотров13 реакций28 пересылок

    🍿 Dawn of Generative Cinematography 🍿 🟩 The Odyssey by Christopher Nolan was shot entirely on IMAX 70mm. It feels almost romantic: massive cameras, film stock, premium lenses, and an obsessive pursuit of the highest possible quality at the moment of capture. 👉 Meanwhile, AI research is heading in the exact opposite direction. 🟩 A pre-print paper released today, "Camera Anything", demonstrates something that sounded like science fiction just a few years ago: you film a scene once... and then you can virtually reposition the camera anywhere. 👉More: https://t.ly/g-qUh 👉Paper arxiv.org/pdf/2607.24591 👉Project yixuanli98.github.io/cameraanything/ 👉Repo github.com/yixuanli98/CameraAnything

  • 29 июл.2 770 просмотров10 реакций27 пересылок

    🔥Decoder-only Any-to-Any Model🔥 👉MODUS unifies any-to-any multimodal generation with one decoder, two experts, and zero task heads. Impressive work. Repo under Apache💙 👉Review https://t.ly/-2QKT 👉Paper https://lnkd.in/dhfBQGhB 👉Project https://lnkd.in/dPD_ECXk 👉Repo https://lnkd.in/dbDHw24u

  • 30 июл.2 668 просмотров15 реакций26 пересылок

    🐠Dual-branch ID-Tracking🐠 👉TIDE: tracking dense, homogeneous targets, providing a scalable dual-branch design to accommodate diverse hardware constraints. MIT license💙 👉Review https://t.ly/WEDeY 👉Paper https://arxiv.org/pdf/2607.26412 👉Project https://vranlee.github.io/TIDE/ 👉Repo https://github.com/vranlee/TIDE

  • 27 июл.2 544 просмотров5 реакций31 пересылок

    💄MagicMakeup Transfer💄 👉Makeup-transfer applies the reference makeup to the source face while preserving the source identity. Authors: Zhejiang University & vivo BlueImage Lab. Repo for non commercial💙 👉Review https://t.ly/JYpCr 👉Paper harxiv.org/pdf/2607.20924 👉Project vivocameraresearch.github.io/magicmakeup 👉Repo github.com/vivoCameraResearch/Magic-Makeup

  • 28 июл.2 512 просмотров9 реакций31 пересылок

    🔎MicroZoom at Extreme Scale🔎 👉MicroZoom by UWA synthesizes gigapixel-resolution images grounded in consumer-grade microscope close-ups at magnification levels up to 350×. Impressive. Repo under MIT💙 👉Review https://t.ly/hgJD7 👉Paper https://arxiv.org/pdf/2607.24729 👉Project https://microzoom-sr.github.io/ 👉Repo github.com/MicroZoom-SR/MicroZoom-SR.github.io

  • 26 июл. просмотров

    без подписи