tgindex
Henok | Neural Nets

Henok | Neural Nets

Статистика

Group: https://t.me/neural_netss_chat

Последний пост
17 июн.
Последнее чтение
15 авг.
Постов за неделю
0
Всего постов
22
Тип
открытый
Язык
und
В каталоге с
14 авг.
Подписчики
2 239
+1 за 2 дн.
Сутки
0
0,00%
Неделя
 
Месяц
 
Просмотров на пост
1 033
22 постов
Вовлечённость
46,1%
к подписчикам
Постов в день
0,0
всего 22
Упоминаний
0
каналов
Охват размещения
оценка
1/24сутки в ленте
1/48двое суток
1/72трое суток

Оценка по просмотрам недавних постов: пост набирает почти всё за первые сутки.

Посты

  • 17 июн.1 014211из birhan_nega

    የእግርኳስ ሆርሙዝ ሰርጥ 🔥

  • 16 июн.900175из Dagmawi_Babi

    Introducing Papers API

  • 13 июн.1 180232

    Today in ባህር ዳር

  • 13 июн.1 0112

    Diogenes the Cynic's story is interesting. If I was a member of the parliament I'd have entered with lit lantern😂 https://penelope.uchicago.edu/encyclopaedia_romana/greece/hetairai/diogenes.html

  • без подписи

  • so for world cup, I predicted 2/2 correct predictions 😎, I mean who is stopping me now. I dare you to ask me who the next PM of 🇪🇹 is going to be

  • Diffusion Gemma is really nice. Making me go back to abandoned diffusion based text generation projects https://deepmind.google/models/gemma/diffusiongemma/

  • Diffusion Gemma is really nice. Making me go back to abandoned diffusion based text generation projects https://deepmind.google/models/gemma/diffusiongemma/

  • 9 июн.1 025198

    New 4B translation model from Hasab AI for few Ethiopian langs. It's good to see almost everyone who is working on AI/ML in Ethiopia is releasing something. This will help a lot to make a progress. 👏 Hasab AI Now tag Ethiopian AI Institute to release something too 😁 https://huggingface.co/hasab-ai/YehaTranslate

  • Just watched this, and they are doing cool things. It's also a hard decision to be an open source AI company esp in Ethiopia👏👏👏 I want to see them succeed. This weekend I will check their models, datasets etc and run a few experiments.

  • My weekly GPU usage📊 ~1267.2 kWh According to Gemini this can power an average household refrigerator for 14 to 16 months, or driving a standard electric vehicle (EV) for about 6,000 kms The good thing is the cluster runs on renewable energy so close to zero carbon footprint and recycles the heat back.

  • 6 июн.2 26214

    без подписи

  • 6 июн.2 2602914

    Oh wow

  • 5 июн.2 300189

    Saw a post about ScholarXIV from Babi, so you heard about fake citations right, well ScholarXIV can help with that🔥 But open research problem could be how can you ground the LLMs so they can cite exactly from the paper without altering results, no errors or attach the citation to the wrong sentence etc Maybe methods like on-policy distillation could help here, let a verifier identify where the model’s claim diverges from the source, then train the model to down weight those unsupported claims, wrong source sentence citations etc.

  • So I got stuck making an objective function with an anti collapse loss so the embedding space doesn’t collapse into a useless low dimensional subspace 😞 maybe venting here helps lol, tips are also welcomed

  • 4 июн.902171

    But we had a retweet from him😁, not as good as Arsenal's win but a win is a win

  • 4 июн.764101

    I've been a fan of Sasha for quite a while. I even tried emailing him and applied to be advised by him in 2022 when he was at Cornell Tech. I didn't get a reply on that email 😂. He got great students, he now works at Cursor, if you follow this channel I also post some of the puzzle he made like the GPU puzzle etc. This was a post from Dwarksh yesterday. The point is, don't be shy just email any researcher, start today, who cares. You either get better mentors, or watch that person get famous and flex saying I've emailed that guy before 😂 Recently met @srush_nlp and he started giving me an impromptu lecture on how targeted on-policy self-distillation works. I asked him if I could record it on my iPhone. The basic idea is this: if the model made a mistake at some point in the rollout (for example, calling a tool that doesn't exist), we want to discourage this specific error, but we don't want to just learn from the final reward, because it's a very noisy signal spread out over the whole trajectory. So we have another model read this trajectory and figure where the error was made. It simply inserts some hint tokens to the part of the trajectory right above where the mistake was made. Now with these injected hint tokens, have the model run a forward pass. You're not having to regenerate a new rollout - aka no new decode required. The hint causes the model to assign lower probabilities to the error tokens. You then trains the original model to match these new probabilities, teaching it to downweight that specific mistake. https://x.com/dwarkesh_sp/status/2062353335529935114?s=20

  • без подписи

  • без подписи

  • без подписи