Artificial Intelligence || DL
СтатистикаChannel for who have a passion for - * Artificial Intelligence * Machine Learning * Deep Learning * Data Science * Computer vision * LLMs and NLP Admin: @idrokdev
- Последний пост
- 3 авг.
- Последнее чтение
- 13 авг.
- Постов за неделю
- 0
- Всего постов
- 22
- Тип
- открытый
- Язык
- und
- В каталоге с
- 13 авг.
- 1/24сутки в ленте
- 140
- 1/48двое суток
- 160
- 1/72трое суток
- 173
Оценка по просмотрам недавних постов: пост набирает почти всё за первые сутки.
Посты
🐠Dual-branch ID-Tracking🐠 👉TIDE: tracking dense, homogeneous targets, providing a scalable dual-branch design to accommodate diverse hardware constraints. MIT license💙 👉Review https://t.ly/WEDeY 👉Paper https://arxiv.org/pdf/2607.26412 👉Project https://vranlee.github.io/TIDE/ 👉Repo https://github.com/vranlee/TIDE
🍿 Dawn of Generative Cinematography 🍿 🟩 #TheOdyssey by Christopher Nolan was shot entirely on IMAX 70mm. It feels almost romantic: massive cameras, film stock, premium lenses, and an obsessive pursuit of the highest possible quality at the moment of capture. 👉 Meanwhile, #AI research is heading in the exact opposite direction. 🟩 A pre-print paper released today, "Camera Anything", demonstrates something that sounded like science fiction just a few years ago: you film a scene once... and then you can virtually reposition the camera anywhere. 👉More https://t.ly/Kd7RV 👉Paper arxiv.org/pdf/2607.24591 👉Project yixuanli98.github.io/cameraanything/ 👉Repo github.com/yixuanli98/CameraAnything
💢Unified Video Dense Prediction💢 👉UniD predicts: depth, surface normals, semantic segmentation, boundaries, human parts, albedo, shading, and materials. Code TBR💙 👉Review https://t.ly/oo7et 👉Paper https://arxiv.org/pdf/2607.21592 👉Project https://unid-video.github.io/ 👉Repo https://github.com/YihongSun/UniD
👉Not a render. Not a concept. This is GENE.01 by Generative Bionics, the Italians coolest scaleup strikes back: in just six months, they turned GENE.01 into a fully functional humanoid platform that can walk, sense and interact. 👉Full-body multimodal skin perceives touch, proximity, force and temperature, bringing Physical AI closer to safe and natural collaboration with people. 👉More: https://t.ly/F3I3A
🏯SOTA Music-to-Dance Gen🏯 👉The Tongyi Lab unveils Wan-Dancer, a novel stable minute-scale synthesis at 720p/30fps across five dance genres. Impressive results, new SOTA on long clip by a large margin. Repo under Apache 2.0💙 👉Review https://t.ly/AKY5j 👉Paper https://lnkd.in/d_xA7dwb 👉Project https://lnkd.in/dzfnw2h4 👉Repo https://lnkd.in/d-Zj_cTf
🌈FlowWAM: flow->action prediction🌈 👉FlowWAM is a novel dual-stream diffusion framework that adopts optical flow as a unified, video-native action representation. Repo under Apache💙 👉Review https://t.ly/FmutT 👉Paper https://arxiv.org/abs/2607.13017 👉Project https://flow-wam.github.io/ 👉Repo github.com/YixiangChen515/FlowWAM
🔥ZipDepth: Depth on Any Device🔥 👉ZipDepth from UniBO is a super-compact monocular depth network by combining an efficient reparameterizable encoder-decoder with large-scale knowledge distillation from a foundation model. Repo under MIT💙 👉Review https://t.ly/qYrLZ 👉Paper https://arxiv.org/pdf/2607.08771 👉Project https://zipdepth.github.io/ 👉Repo https://github.com/fabiotosi92/ZipDepth
🌔Foundation Global SFM🌔 👉Glob3R is a global SfM-style reconstruction built on 3D foundation models. key idea: explicitly optimize feed-forward geometric predictions. Repo TBA💙 👉Review https://t.ly/Z_4C7 👉Paper https://arxiv.org/pdf/2607.09225 👉Project https://junyuandeng.github.io/Glob3r/ 👉Repo TBA
💋SAM-MT: Real-Time Multi-Target VOS💋 👉Fudan & Shangai unveil SAM-MT, an efficient interactive multi-target video segmentation framework that maintains near-single-object efficiency (FPS/VRAM) as target count increases, while maintaining robust video segmentation performance. Repo available💙 👉Review https://t.ly/Z_4C7 👉Paper https://lnkd.in/dvS-iyBD 👉Project https://lnkd.in/daQ8na8T 👉Repo https://lnkd.in/dgbX2tZv
🍀OctoSense: Open Sensing🍀 👉OctoSense is an open-source sensor platform with stereo RGB and event cameras, LiDAR, a thermal camera, an inertial measurement unit, RTK-corrected global positioning system, and proprioception. 👉Review https://t.ly/oFN8L 👉Paper https://lnkd.in/dM3zpyju 👉Project https://lnkd.in/ddrQ3uJ6 👉Repo https://lnkd.in/dhSDjSfG
🔊VolHuMe - Volumetric Human Meshes🔊 👉VolHuMe (H/T @Martinella_94) is a novel, high-resolution large-scale dataset of volumetric human meshes with complete 4D GT: multi-view RGB-D, textured meshes, dense point clouds, normal maps, rigged assets, garment segmentation, and SMPL-X fittings in one dataset. Insane💙 👉Review https://t.ly/b5vxy 👉Paper https://arxiv.org/pdf/2606.23062 👉Project giuli13.github.io/volhume-website/# 👉Repo TBA soon
Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal Prior 👉Paper https://moygcc.github.io/vid2avatar-pro/static/CVPR2025_Vid2Avatar_Pro.pdf
🕷️Human Universal Grasping🕷️ 👉HUG is a flow-matching model that generates diverse human grasps for any user-specified object in a single RGB-D image captured from a stereo camera. 👉Review https://t.ly/VG1Eu 👉Paper https://arxiv.org/pdf/2606.17054 👉Repo https://github.com/KevinyWu/hug 👉Project https://grasping.io/
🔍 Nvidia Locate Anything 🔍 👉Diverse localization tasks under a unified vision-language model, including document understanding, GUI grounding, dense detection, and OCR. Repo released💙 👉Review https://t.ly/PvwFo 👉Paper https://lnkd.in/dWfNpzPZ 👉Project https://lnkd.in/dM89BX-8 👉Repo https://lnkd.in/dC4KCQSM
видео или голосовое, без подписи
🪔Latent Decoding Pixel Diffusion🪔 👉PiD by Nvidia is a plug-and-play diffusion decoder that replaces VAE/RAE decoders, turning latent representations directly into super-resolved pixels in a single pass. Repo under Apache 2.0💙 👉Review https://t.ly/y19mA 👉Paper https://lnkd.in/duVC25C2 👉Project https://lnkd.in/dW6TkzCB 👉Repo https://lnkd.in/dnGdgKRr
🍒Count Anything, Any Granularity🍒 👉Open-world counting as multi-grained counting, where visual exemplars specify target appearance and fine-grained text specifies the intended semantic granularity across five explicit levels. Repo/Data under Apache💙 👉Review https://t.ly/nqz80 👉Paper https://lnkd.in/dp7khTRU 👉Project https://lnkd.in/d_jfX_Yn 👉Repo https://lnkd.in/dkTRGZkG 👉Data https://lnkd.in/dB83jRyT
🦄Unified Correspondence Transformer🦄 👉UniCorrn is the first correspondence model with shared weights that unifies 2D-2D, 2D-3D, and 3D-3D geometric matching with a transformer. CC BY-NC-SA 4.0💙 👉Review https://t.ly/2OBdq 👉Paper https://arxiv.org/pdf/2605.04044 👉Project https://neu-vi.github.io/UniCorrn/ 👉Repo https://github.com/neu-vi/UniCorrn
🪝Syn4D: Multiview Synthetic 4D Dataset🪝 👉Syn4D is novel multi-view synthetic dataset of dynamic scenes that includes ground-truth camera motion, depth maps, dense tracking, and parametric human pose annotations💙 👉Review https://t.ly/SL1mk 👉Paper https://arxiv.org/pdf/2605.05207 👉Project https://jzr99.github.io/Syn4D/ 👉Repo https://github.com/jzr99/Syn4D 👉Data huggingface.co/datasets/Syn4D/Syn4D_RGBD/tree/main
🧘♀️Holistic Shot Boundary Detection🧘♀️ 👉OmniShotCut detects shot changes of the video in diverse sources (anime, vlog, game, shorts, sports, screen recording, etc.), and recognize Sudden Jump and Transitions (dissolve, fade, wipe, etc.) by proposing a Shot-Query-based Video Transformer. Repo, demo & benchmark💙 👉Review https://t.ly/sTi7N 👉Paper https://arxiv.org/pdf/2604.24762 👉Project uva-computer-vision-lab.github.io/OmniShotCut_website/ 👉Repo github.com/UVA-Computer-Vision-Lab/OmniShotCut