tgindex
Ontology Official Announcement

Ontology Official Announcement

Статистика
Последний пост
29 июл.
Последнее чтение
15:01
Постов за неделю
0
Всего постов
24
Тип
открытый
Язык
английский
Категория
Криптовалюты (по похожим)
В каталоге с
13 авг.
Подписчики
2 945
−2 за 5 дн.
Сутки
−5
−0,17%
Неделя
 
Месяц
 
Просмотров на пост
526
24 постов
Вовлечённость
17,9%
к подписчикам
Постов в день
0,0
всего 24
Упоминаний
0
каналов
Охват размещения
оценка
1/24сутки в ленте
425
1/48двое суток
487
1/72трое суток
525

Оценка по просмотрам недавних постов: пост набирает почти всё за первые сутки.

Посты

  • 📌 Ontology MainNet upgrade to v3.1.2 Ontology will perform a scheduled MainNet upgrade to v3.1.2 at block height 20,800,000. What is in it 🔹 PUSH0 (EIP-3855) places zero on the stack in one byte and 2 gas, instead of two bytes and 3 gas. 🔹 BASEFEE (EIP-3198) lets a contract read the current base fee directly on-chain for 2 gas, removing the need for an external data source. 🔹 MCOPY (EIP-5656) copies memory in a single instruction. Copying 256 bytes drops from at least 96 gas to 27, which matters for encoding, byte handling and cryptographic work. 🔹 Transient storage, TSTORE and TLOAD (EIP-1153) adds low-cost state that lasts for one transaction only, at 100 gas each. Node operators All node operators should upgrade to v3.1.2 as soon as possible and before block 20,800,000 is reached. Release: https://github.com/ontio/ontology/releases/tag/v3.1.2 Full breakdown 👉 https://ont.io/news/ontology-mainnet-v3-1-2-four-ethereum-opcodes-arrive-on-the-ontology-evm/

  • 📌 Ontology at Eight: verified human data for the AI economy Eight years ago the Ontology MainNet went live, and it has run without interruption ever since. But our eighth anniversary is not really about looking back. It is about what all of it was building toward. AI is only as good as the data it learns from, and the industry is moving away from scraped content toward high-quality human data: consent-based, and provably created by a real person. The problem is supply. The people who create data rarely share in its value, while big tech has earned more than $1.3 trillion from user-generated data. That is the gap we have spent eight years preparing to fill. ONTO Wallet stays a multi-chain Web3 wallet and is now adding an identity and verified human data platform: you own the data you create, build a verified profile, and earn rewards by contributing it on your own terms. Read the full piece 👉 https://ont.io/news/ontology-verified-human-data/

  • 📌 When human oversight becomes a compliance requirement "We had humans in the loop" is a description of a process. "Prove it" is a demand for evidence, and evidence has properties good intentions do not. It has to name specific people, show they were distinct real humans rather than sybils or one-shot contractors, show their judgement held up over time, and make every contribution attributable, timestamped and tamper-evident. The ground is already moving. The EU AI Act requires human oversight for high-risk AI under Article 14, and the frontier labs themselves are warning that recursive self-improvement could quietly drop the human from the loop. The moment oversight has to satisfy an auditor, attestation ("trust us, qualified humans reviewed this") is not enough. You need provenance: a record someone who does not trust you can check. Run the self-check 👉 https://ont.io/news/human-oversight-documentation/

  • без подписи

  • 📌 New: when the judge shares the blind spot Roll two fair dice. You are told at least one is a six. The probability that both are sixes is not 1 in 6, it is 1 in 11: the clue removes every outcome with no six, leaving 11 equally likely cases, one of which is the double six. The fast answer assumes an independence the clue already broke. Avena et al. tested eight state-of-the-art models on problems built exactly to trigger that shortcut. On the counterintuitive items the models failed consistently and predictably, and chain-of-thought did not reliably rescue them. The trouble is that those same models now do the grading. When an LLM-as-judge carries a reasoning blind spot, a reward model trained on its preferences inherits it, and model-on-model agreement certifies consensus rather than correctness. The only check that does not share the failure mode is human ground truth you can verify: evaluators whose reasoning consistency is measured and tracked over time, carried on a stable identity with signed contributions. W3C Decentralized Identifiers, W3C Verifiable Credentials, W3C Bitstring Status Lists. ONT ID and ONTO Wallet are the substrate. Day 1 of Ontology Roundup, Issue 04. Try the dice trap 👉 https://ont.io/news/llm-as-judge-blind-spots/

  • 🎲 We're running a little experiment today, and we want you in it. One dice question. One trap almost everyone falls for, the AI models included. The poll is live on X right now 👉 https://x.com/OntologyNetwork/status/2066416847499485589?s=20 Vote, argue it out in the replies, show your working, and tag a friend who reckons they're good at probability. The more wrong answers, the better the point we're making. No Googling. No AI. That is cheating, and it gets this one wrong anyway. Reveal and the full piece this afternoon. 👀

  • без подписи

  • без подписи

  • без подписи

  • без подписи

  • без подписи

  • A checklist today, because the SFT-vs-RL argument is eating attention that belongs one layer up. Whichever recipe wins for reasoning models, both consume step-level human evaluation, and almost nobody can defend theirs. Five questions sort the defensible pipelines from the rest: who made each judgement, does it carry its rubric, would you notice a drifting evaluator, can experts prove credentials without exposing identity, and does revocation actually propagate. Fewer than three yes answers means the pipeline, not the recipe, is the binding constraint.

  • 📌 New: Evaluator-backed benchmarking, after the MLE-Bench moment MLE-Bench has been quietly contested across r/MachineLearning over the last week. The skepticism is not really about any single metric inside the benchmark; it is about whether a static benchmark structure can survive sustained adversarial attention from teams with economic incentive to game it. The standard answer (better methodology, rotating held-out sets, broader task coverage) is real and partial. None of it fixes the structural problem: the benchmark as an artefact is a fixed target. Evaluator-backed benchmarking is the structural counter. Every judgement contributing to a published benchmark statistic traces back to a stable evaluator identity (a W3C DID the evaluator controls), a signed verifiable credential carrying the rubric version and expertise attestations, longitudinal consistency credentials, and a status trail for revocations. The benchmark stops being a number the publisher asks the field to trust. It becomes an artefact any third party can audit at the judgement layer. Issue 02 Monday made this argument at the policy-and-research level with the METR teardown. MLE-Bench is the same warning shot moved one level closer to the user-facing capability claim. The first publishers to ship evaluator-backed benchmarking will be the ones whose results survive the next round of teardowns. This is Day 2 of Ontology Roundup, Issue 03. Read it 👉 https://ont.io/news/evaluator-backed-benchmarking/

  • 📌 New: Reward models need reward-model QA The recent LongTraceRL work made one thing unavoidable: sparse outcome signals are not enough at the reasoning-trace layer. The field has to evaluate intermediate reasoning steps. Step-level evaluation is a substantially different operation than outcome evaluation: judgements are finer-grained, cognitive load on the evaluator is higher, and the noise floor on any individual rating is correspondingly worse. A reward model trained on step-level data is more sensitive to evaluator quality than the outcome-level reward models the field is used to. Sloppy step-level judgement does not just add noise; it miscalibrates the reward model in structured ways the team training the model may not be measuring. Reward-model QA is the missing layer that turns step-level preference data into trustable training signal. The standards stack is the same one Issue 02 set out: W3C Decentralized Identifiers anchor stable evaluator identity, W3C Verifiable Credentials carry signed step-level contributions with rubric versioning, W3C Bitstring Status Lists handle revocation. With those in place, the reward-model team can defend who made each judgement, under what methodology, with what calibration history. This is Day 1 of Ontology Roundup, Issue 03. Read it 👉 https://ont.io/news/reward-model-qa-longtracerl/

  • 📌 New: The evaluator uniqueness primitive: from sybil resistance to agent evaluation Closing piece for Issue 02. The week opened with the METR teardown and traced the credibility-event pattern through preference data integrity (Tuesday) and longitudinal evaluation (Wednesday). This piece folds together the two threads still open: chronic sybil contamination in preference-data marketplaces, and the agent decision evaluation vacuum that has not yet crystallised into a named problem. Both are solved by the same primitive. Selective disclosure (W3C Verifiable Credentials 2.0 + IETF RFC 9901 SD-JWT) lets a credentialed issuer attest that an evaluator is one unique person, certified by a trust framework, without disclosing identity or demographics. The reward-model team gets the uniqueness guarantee. The evaluator gets privacy. Neither has to compromise. The same mechanic becomes the ground-truth layer for agent decision evaluation when that problem crystallises later this year. The primitive that closes all five Issue 02 threads (benchmark provenance, preference data integrity, longitudinal evaluation, sybil resistance, agent decision evaluation) is human judgement with verifiable uniqueness. The standards work has been done. ONT ID and ONTO Wallet are the substrate. This is Day 5 of Ontology Roundup, Issue 02. Closing piece. Read it 👉 https://ont.io/news/evaluator-uniqueness-closer/

  • 📌 New: Continuous training needs continuous evaluators Deployed models do not sit still anymore. Retrained, fine-tuned, instruction-extended, behaviourally patched on a cadence measured in weeks, sometimes in days. Last week's Prism paper (Tang et al., arXiv 2605.26110) treats multimodal continual instruction tuning as the deployed reality and flags that the field is hindered by severe engineering bottlenecks. The bottlenecks on the model side are well-defined. The bottlenecks on the evaluation side are larger and quieter. A snapshot evaluator pool against a continually retrained model is the slow version of a contaminated reward dataset. The published delta between version N and N+1 is the sum of two things: actual model behaviour change and cohort composition change. Most teams cannot separate the two. The cost surfaces months later as benchmarks that no longer agree with one another and methodology questions that cannot be resolved without going back to data the pipeline did not keep. Longitudinal evaluation is the property that the evaluator cohort is observable over time, the same way the model is. Stable evaluator identity across batches, signed and timestamped contributions, auditable cohort composition. W3C Decentralized Identifiers, W3C Verifiable Credentials, W3C Bitstring Status Lists. The standards have been mature for years. This is Day 3 of Ontology Roundup, Issue 02. Read it 👉 https://ont.io/news/longitudinal-evaluation/

  • 📌 New: Your reward model is only as good as your preference data Last week's RTDMD paper (Huang et al., arXiv 2605.26108) proposes reward-guided RL for few-step diffusion alignment. It also explicitly acknowledges that aligning distilled models with human preferences remains challenging. The framework solves a downstream optimisation problem; the upstream supply of preference signal still does what it has always done, which is determine the ceiling on everything built on top of it. A reward model trained on inconsistent, sybil-contaminated, or methodologically opaque preferences encodes those defects, and distillation propagates them faster at lower latency. Efficiency at the model layer does not fix a quality problem at the judgement layer. It amplifies it. Preference data integrity is the property that every preference judgement can be traced back to a stable evaluator identity, a signed rubric at the version that applied, a verifiable record of the evaluator's credentials, and a status trail showing what has been revoked or superseded. The standards stack is mature: W3C Decentralized Identifiers, W3C Verifiable Credentials, W3C Bitstring Status Lists. The teams that build distillation pipelines on top of preference data with verifiable integrity will be the ones whose aligned models actually do what their alignment claims say they do. This is Day 2 of Ontology Roundup, Issue 02. Read it 👉 https://ont.io/news/preference-data-integrity-the-variable-distillation-hides/

  • 📌 New: When benchmarks break: the case for traceable evaluator provenance The METR time-horizons graph, cited everywhere from policy briefings to capability roundups, has been publicly contested. A detailed teardown documents what one critic described as numerous severe errors. The graph had become a load-bearing reference for an industry whose habit is to cite it and move on. This is not about METR specifically. It is the latest, loudest instance of a category: benchmarks that have no traceable evaluator provenance behind them. When the underlying judgement chain is opaque, any methodology question becomes an unfalsifiable argument. Trust collapses to authority, and authority is exactly what every external observer was already sceptical of. Evaluator provenance is the verifiable chain from "this person made this judgement at this time, using this rubric" through to the aggregate statistic. The standards stack has been mature for years: W3C Decentralized Identifiers, W3C Verifiable Credentials, W3C Bitstring Status Lists. With it in place, who made each judgement, what methodology was attested, and what has happened to the credential since all become observable to anyone outside the publishing organisation. This is Day 1 of Ontology Roundup, Issue 02. Read it 👉 https://ont.io/news/evaluator-provenance-metr/

  • 🎮 Ontology x PALZ Quiz Time! This Friday's Discord Community Quiz has a special PALZ round, with $ONG prizes on the line. 📅 Friday, May 29 🕘 9AM UTC 📍 Ontology Discord 💡 Tip: spend a bit of time in the PALZ game beforehand. Knowing your way around the islands and your favourite PAL will give you the edge when the PALZ round drops. See you there. 🐾 👉 Join the quiz: https://t.co/3aa6WNHXgi 🎮 Play PALZ: https://www.palzgame.com

  • 📌 New: Signed content for a world where platforms are AI AI-mediated communication systems measurably shift the opinions of the groups they serve. Polish, suggest, summarise, rewrite. Each tap nudges. The aggregate shifts. "Did this person say this thing" is becoming a real question, not because identities are forged, but because the path from "what the human meant" to "what arrived on the platform" now routinely runs through a model. Durable content provenance requires three layers working together: C2PA for structured manifests (who, when, what tools, what edits, AI involvement), blockchain for platform-independent anchoring (manifests survive arbitrary hops because the cryptographic anchor is not attached to the file), and decentralised identity to bind the signer (a DID outlives any platform that issued it). C2PA alone is necessary but fragile. Blockchain alone does not capture the structured provenance C2PA provides. Decentralised identity alone has no content to attest to. Pair the three and the question of authorship becomes tractable again. This is Day 5 of the Ontology Roundup, Issue 01. Read it 👉 https://ont.io/news/content-provenance-ai-platforms/