CS
coloredskyscore™
Back to dashboard
Governing document

Project Constitution

The following is the constitution for the coloredsky score project — the opportunity it answers, its mission, the principles it holds itself to, and the test any proposed change to the methodology has to pass.

Article IThe Opportunity

Blockchain claims are abundant. Verified fundamentals are not. Every chain claims to be faster, safer, more decentralized, or more institution-ready than the last, and there are remarkably few reliable ways to actually determine how true any of that is — or to hold one chain's claims against another's on equal terms.

The public discourse around any given chain has moved through several generations — forums, then Reddit, then YouTube, then X — and gotten louder at each step without getting more rigorous. Paid influencers now shape a meaningful share of what circulates as analysis, and their incentives are rarely disclosed alongside their opinions. A confident take and a correct one are not distinguishable from the outside.

Primary sources exist but don't close the gap. Whitepapers describe design intent, not verified operating reality, and reading one well enough to judge it requires exactly the technical background most people evaluating a chain don't have. Two whitepapers, read in good faith, don't tell you which chain is actually stronger — they tell you which project wrote the more persuasive document.

First-hand experimentation is the most honest signal available and still isn't sufficient. Using a chain yourself is real evidence, but it measures user experience — speed, cost, app quality — not the underlying fundamentals: regulatory standing, security history, supply structure, decentralization, confidentiality capability. A chain can feel excellent to use and still be structurally weak, or the reverse, and hands-on testing alone can't tell the two apart. It's also not scalable: it works for the two or three chains someone actually uses, not for a landscape of dozens.

The compounding problem is that none of these sources evaluate chains against the same criteria. Every comparison — a forum thread, an influencer thread, a personal trial — is implicitly apples-to-oranges, even when everyone involved is acting in good faith. There is no equivalent of a credit rating for blockchain fundamentals: a fixed, disclosed methodology applied identically to every chain, so that a score means the same thing wherever it's attached.

That absence is the opportunity. A scoring system that removes the guesswork — that lets someone without deep technical expertise trust a number instead of personally decoding whitepapers, docs, and influencer takes to arrive at an opinion.

Thus, coloredsky score was created.

Article IIThe Mission

Score blockchains against fixed, disclosed, evidence-based criteria — the same criteria applied to every chain — so that fundamental quality can be compared the way a bond or equity rating can be compared.

"Is this chain any good?" is not actually a single question, so this system doesn't try to answer it with a single number. It's broken into two distinct lenses. The first is TradFi — short for Traditional Finance, meaning banks, asset managers, pension funds, and the rest of the regulated institutional investing world — and asks whether a chain is structurally ready for that kind of institutional capital. The second asks whether an autonomous AI agent could actually operate on the chain unattended. A chain can score excellently on one lens and only mediocre on the other, and that divergence — not either raw score on its own — is frequently the most useful thing the system produces.

Working thesis, stated honestly as a thesis and not as a scored input: if the methodology is sound, chains that score higher should tend to attract more capital than chains that score lower, despite markets being imperfect and sometimes irrational in the short run. The scoring system is not built to chase that outcome — narrative, sentiment, and price are explicitly out of scope — but validating or falsifying that thesis over time is a real reason this project exists.

Success looks like: the chain's scored numbers being something a retail investor trusts more than their own read of a whitepaper or an influencer thread, something a TradFi allocator recognizes as rigorous enough to replace their own research team's diligence, and something an AI agent can query directly as infrastructure rather than a human having to interpret it on the agent's behalf.

What this is not

To make this scoring system effective and scalable, we abide by the following principles.

Article IIIPrinciples

Section 1 — Evidentiary Standard

1.1Score reality, not aspiration.
Only what is verifiably live today counts. Roadmaps, testnets, and announcements earn zero credit until they ship.
1.2Narrative, sentiment, and price never move a score.
If a proposed change would only be justified by "the market would like this" or "this chain gets talked about more," it does not belong in the methodology.

Section 2 — Framework Integrity

2.1Framework integrity outweighs any single chain's score.
A rule is not bent to produce a flattering — or unflattering — result for the chain in front of you. Decisions get made and then applied to whichever chain the chips fall on.
2.2Cross-chain comparability is the product.
A change that makes one chain's story read better but breaks the ability to line every chain up on the same ruler is a net loss, even if it's individually correct.

Section 3 — Structural Fairness

3.1Withheld vs. unknowable are different, and the direction of the penalty depends on which one it is.
If someone could publish a fact and hasn't, score the unfavorable end. If nobody can know it by design, don't score it at all — don't penalize architecture.
3.2A structural fact can legitimately produce different consequences in different places
A credit here, a debit there, a gate call and a category note — as long as each place is scoring a genuinely distinct manifestation, never the identical fact twice.

Section 4 — Gates, Categories, and Divergence

4.1Gates are disqualifying checks, evaluated first, and they set a ceiling — not a score.
Each framework has a small, fixed set of gates covering the handful of properties severe enough that failing one should limit how high a chain can score no matter how strong everything else is. Each gate is judged Pass, Partial, or Fail before any category is scored. A full Pass places no ceiling on the final number. A Partial or Fail attaches a cap — a maximum the chain's final published score cannot exceed, applied after categories are calculated, using whichever gate's cap is lowest if more than one is capped. Gates are a floor, not a guarantee: passing every gate means nothing disqualifying is present, not that the chain scores well. A chain can clear every gate cleanly and still land in a mediocre tier, because categories — not gates — are what determine how good a chain actually is.
4.2Categories are graded strength, weighted and summed, and they're what actually produces the score.
Where a gate asks a narrow disqualifying question with three possible answers, categories are the opposite: a wider set of specific fundamentals, each scored on a continuous scale reflecting how strong the chain genuinely is on that one dimension, each carrying its own weight toward the total, summed into a single raw weighted score before any gate cap is applied. Categories are where real differentiation between chains happens — two chains that both cleanly pass every gate can still land in very different tiers, because the gates only screened out disqualifying weaknesses and said nothing about relative strength. A gate answers "is anything disqualifying present." Categories answer "how good is this, specifically, compared to every other chain scored the same way."
4.3Divergence is signal, not noise.
The most useful outputs this system produces are gaps, not single numbers. TradFi-vs-AI divergence is one: the same chain scoring well on one axis and poorly on the other. The other is external to scoring entirely — a chain's fundamentals score becomes a baseline once published, and a token's market price can sit well above or below what that baseline implies. That gap is never fed back into the score itself (price is explicitly out of scope per §1.2), but it's a legitimate and useful thing to point at afterward: a chain trading rich or cheap relative to its own fundamentals. Nothing should smooth either gap away for the sake of a tidier number.

Article IVScoring Methodology Change Management

Before a new idea becomes a methodology consideration — rather than a passing thought, or nothing — it should be able to answer these, roughly in order.

01 — DOES IT CHANGE WHAT WE MEASURE, OR ONLY HOW CONFIDENT WE ARE?
A genuinely new signal is different in kind from a data-quality fix. Both are legitimate, but they cost differently and should be logged as different kinds of items.
02 — IS IT STRUCTURAL, OR NARRATIVE WEARING A STRUCTURAL COSTUME?
"This chain's community is excited about X" is never an item. "This chain shipped X to mainnet and no existing category or principle prices it" might be. If the honest one-sentence justification leans on sentiment, price, or vibes, it fails here regardless of how it's phrased.
03 — DOES PRECEDENT ALREADY ANSWER IT?
Check the Principles and closed items first. A real question raised, matched to an existing ruling, and resolved in the same session it came up gets recorded as answered rather than left open. Reinventing an answered question is waste; the fix is citing precedent, not opening a new item.
04 — WHAT DOES ADOPTING IT ACTUALLY COST, AND DOES THAT COST MATCH ITS VALUE?
Every open item should be honestly labeled: rubric language only, one category's re-research, a full board re-research, or a gate-level change — which can move a binding cap and therefore a published score by more than any category ever could. A gate-level idea gets more scrutiny before it's even logged, not just before it's adopted.
05 — DOES IT PRESERVE COMPARABILITY, OR QUIETLY BREAK THE RULER FOR CHAINS ALREADY SCORED?
Any change that would move published numbers needs an honest, explicit answer to what happens to the chains already scored before it's adopted — recalculating the whole board immediately, deliberately leaving prior scores as historical snapshots, or applying the change forward-only from here. What isn't acceptable is adopting a change and leaving that question unanswered. Silently drifting is not an option.
06 — IS NOW ACTUALLY THE TIME?
Some ideas are correct but premature — held deliberately until a real-world trigger condition is actually met, rather than acted on early because it seems inevitable. A good idea logged for later is a success, not a failure to act.
If an idea survives all six honestly, it earns a place in the methodology backlog, tagged with which cost tier it falls into. If it fails on narrative or is already resolved by precedent, it doesn't get logged as open — it gets logged as decided, so the question doesn't have to be re-derived from nothing next time it comes up.

Maintenance

This document is meant to be stable. It should not need to change often, if ever — the day-to-day evolution of the methodology happens in the internal frameworks and documentation this Constitution sits above, not here. A change to this document is warranted only by something at the level of the mission itself shifting, such as a major change in the blockchain industry this Constitution did not anticipate. When that happens, it's a constitutional change, and it should be discussed and settled with the same deliberateness the frameworks apply to their own principles — not edited casually.