Content Dimensions
Report on content hierarchy, article metadata and reader context as first class dimensions.
Overview
Availability
| Capability | Web | iOS / tvOS | Android / Android TV | Tizen |
|---|---|---|---|---|
setDimension() / setDimensions() | ✅ | ✅ | ✅ | ✅ |
getVisitSequence() / getEngagementLevel() | ✅ | ✅ | ✅ | ✅ |
newsletter from the opening link (handleOpenURL()) | ✅ | ✅ | ✅ | — (left undeclared) |
| Automatic extraction (hierarchy, JSON-LD, word and media counts) | ✅ | — | — | — |
pushNavigation() | ✅ | ✅ | ✅ | ✅ |
The automatic extraction is web only. It reads the URL path, the page’s JSON-LD and the DOM, none of which exists in a native application. On the native SDKs every content dimension is pushed explicitly by your app, and a dimension you do not push simply keeps its default value.
pageUrl and title are not dimensions on the native SDKs: they are already the
url: and title: arguments of pushPageView. assetUrl is settable there, because
nothing derives it for you.
Beyond the 10 custom data slots, Alke Analytics ships a set of named dimensions that model editorial content and reader context directly. Most of them are filled in automatically by the web collector; every one of them can be overridden by your site.
| Dimension | Filled by | Description |
|---|---|---|
level1, level2, level3 | Automatic | Content hierarchy derived from the URL path |
schemaType | Automatic | Schema.org @type of the page |
datePublished, dateModified | Automatic | Publication and update dates |
isAccessibleForFree | Automatic | Whether the article is free to read — three states: declared free, declared paywalled, or not declared |
wordCount, mediaCount | Automatic | Size of the article body |
coverImage | Automatic | Cover image URL |
loggedIn | Your site | Whether the reader is authenticated |
subscriptionStatus | Your site | Subscription state of the reader |
goal | Your site | Conversion reached on the page |
abtest | Your site | Identifier of the running experiment |
pageGroup | Your site | Free-form grouping of the page — no automatic value |
newsletter, visitSequence | Automatic | Reader loyalty, from the visit cookie. Not settable — read them back with the getters |
Content Hierarchy
The collector splits the URL path on / and drops the empty segments. The last
segment is treated as the article slug and dropped only when the page is a piece of
content — that is, when the schema.org @type of the selected node (see Article
Metadata) is a content type. With no structured data, or with a
generic page wrapper such as WebPage or CollectionPage, no segment is dropped
and the path is taken as it is. A trailing slash plays no part: it is the default
WordPress permalink, so keying on it dropped the section on an article URL and pushed
the slug into level2/level3 everywhere else.
The remaining segments feed level1, level2 and level3.
| Path | @type of the selected node | level1 | level2 | level3 |
|---|---|---|---|---|
/economy/salaries/my-article-84f46e03 | NewsArticle | economy | salaries | — |
/economy/salaries/my-article-84f46e03 | none, or WebPage | economy | salaries | my-article-84f46e03 |
/economy/salaries/ | CollectionPage | economy | salaries | — |
/a/b/c/d | Article | a | b | c |
/a/b/c/d | none | a | b | c |
/my-article | Article | — | — | — |
/my-article | none | my-article | — | — |
On a page carrying no structured data, level3 can therefore hold a slug. That is
accepted — the column is sized for it.
Segments are decoded, lowercased, stripped of control characters and of any / or
\ a decoded %2F would put back, then cut to 65 characters. A harmonized level you
push yourself goes through the same normalization server-side, so it groups with a
derived one.
If your URLs do not map cleanly onto your editorial taxonomy, push a harmonized version instead — see Overriding dimensions.
Page grouping
pageGroup is a free-form grouping of your own choosing, orthogonal to the three
hierarchy levels: a template name, a franchise, an editorial format. It has no
automatic value — nothing is reported unless you set it:
alkeAnalytics.setDimension('pageGroup', 'longform-template');
Up to 128 characters; an empty string removes it, like null. It describes one page,
so it is cleared at every page boundary — see
what the boundary clears.
Article Metadata
Article metadata is read from the page’s JSON-LD blocks
(<script type="application/ld+json">), CDATA-wrapped ones included. The graph is
walked depth-first through @graph, mainEntity, mainEntityOfPage, hasPart and
itemListElement, up to 6 levels deep — WebPage → mainEntity → NewsArticle is
the nominal shape outside SEO plugins, and stopping at the wrapper would report a page
with no metadata at all.
“That page” is your declared <link rel="canonical"> when there is one, and the
current URL without its query string and fragment otherwise — a canonical URL carries
neither, so comparing against the raw address would fail the identity test below on
every campaign arrival (?utm_source=…, ?fbclid=…).
The node describing the page is then picked in this order:
- The node that claims to be this page — an
@id,urlormainEntityOfPageequal to that URL (fragment and trailing slash ignored), preferring a content type over a page wrapper. Without this, a section page listing articles, or an article followed by a block of recommendations, would report the first listed item’s date and paywall as its own. - Otherwise, the single candidate of its family: the content types first, then the generic page wrappers. Both families are closed lists, given below. Several candidates, none of which claims the URL, means the page lists content rather than being it, so the next family is tried.
- Failing both, nothing is reported: an
Organizationor aWebSitenode describes your site, not the page in front of the reader. There is no fallback onto “the first typed node”.
That ordering matters: SEO plugins emit a WebPage node alongside the article
one, in an order you do not control. Preferring the content node means the same page
extracts the same data whichever way the plugin writes its @graph.
Membership is decided by exact name, not by schema.org inheritance: a type absent from both lists below is never selected, however specific it is. The same content list decides whether the last URL segment is dropped as an article slug — see Content Hierarchy.
Content types (30) — the page is the content
Article, NewsArticle, BlogPosting, LiveBlogPosting, TechArticle,
ScholarlyArticle, SocialMediaPosting, AdvertiserContentArticle,
AnalysisNewsArticle, AskPublicNewsArticle, BackgroundNewsArticle,
OpinionNewsArticle, ReportageNewsArticle, ReviewNewsArticle,
SatiricalArticle, Report, Review, Recipe, HowTo, Course,
VideoObject, AudioObject, PodcastEpisode, Episode, Movie, Book,
RealEstateListing, MediaGallery, ImageGallery, VideoGallery
Page wrappers (11) — the page describes or lists content
WebPage, ItemPage, CollectionPage, ProfilePage, AboutPage,
ContactPage, CheckoutPage, FAQPage, QAPage, SearchResultsPage,
CreativeWork
The schema.org namespace is recognised in its http/https and www. forms, so a
canonical "@type": "https://schema.org/NewsArticle" is read like the compact one. A
schema: CURIE is not resolved — its binding lives in @context.
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "NewsArticle",
"datePublished": "2026-03-04T08:30:00+01:00",
"dateModified": "2026-03-04T10:00:00Z",
"isAccessibleForFree": false,
"wordCount": 812,
"image": "https://cdn.example.org/cover.jpg"
}
</script>
Nothing found means nothing reported: the dimension simply keeps its default value.
Three states, not two
isAccessibleForFree, loggedIn and newsletter are stored as three states —
declared true, declared false, and not declared — rather than as booleans.
isAccessibleForFree is only reported when the page actually declares it. A page
whose paywall could not be read reports nothing and is no longer counted as
paywalled, which is what a boolean whose default was “paid” used to do to an entire
site with no paywall. On loggedIn and newsletter, the same distinction separates
“the site declared false” from “the site does not implement this dimension”.
These forms are accepted for isAccessibleForFree: a JSON boolean, "True"/"False"
in any case, the enumeration IRIs http(s)://schema.org/True|False, the integers
0/1 (and their string forms), a one-element array, and an @value wrapper around
any of these. As a last resort, a hasPart[] declaring isAccessibleForFree: false —
Google’s documented pattern for a partial paywall — counts as a false declaration
for the page. Anything else counts as not declared.
Dates and word count
datePublished and dateModified are read as ISO 8601 only (YYYY-MM-DD, with
or without a time component). Anything else is ignored rather than guessed:
04/03/2026 would otherwise parse into a valid and wrong date. With no offset
declared, the value is read as UTC, so the same article does not carry a different
date per reader.
The JSON-LD wordCount is only read when it is a plain integer: "1,234" is refused,
because a lenient parse stops at the separator and reports 1 — a plausible integer
nothing downstream can catch.
image may be a URL, an ImageObject, an array of either, or an @id reference to a
node declared elsewhere in the graph. A reference is only dereferenced when the target
really is an image, so a mis-wired image pointing at the WebPage node does not
store the page URL as a cover. thumbnailUrl is used as a fallback.
Counting words and media
By default the word count comes from the JSON-LD wordCount, falling back to
counting articleBody. Media are not counted, because nothing indicates where the
article body is in the DOM.
Two consequences worth knowing before you build a report on these two columns:
mediaCountstays at0for every page until you configurecontentSelector. Nothing distinguishes “this article has no media” from “media were never measured”, so treat the column as empty on a property that does not set the selector.wordCountcan come from three different instruments — the DOM, the JSON-LDwordCountdeclared by the CMS, or a count ofarticleBody— which do not share a definition of a word. Comparing averages across properties compares instruments; within one property, configurecontentSelectorso every page is measured the same way.
Configure contentSelector to measure the DOM instead — words from the node’s text
with script, style, noscript and template removed, so an embed or an inline
JSON-LD sitting in the article body is not counted as prose; media from the img,
picture, video, audio and iframe elements it contains:
alkeAnalytics.setConfig({
propertyId: 'your-property-id',
contentSelector: 'article.body'
});
Three kinds of element are left out of the media count on purpose: an img nested in
a picture (same media as its parent), an img/video/audio with no source at all
(a lazy-load placeholder), and a cross-origin iframe — in an article body that is
more often an ad slot, a comment widget or a signup form than a media.
If the selector matches nothing, the collector falls back to the JSON-LD as above and logs a warning.
mediaCount stays at 0 on every page and wordCount comes from
whatever the CMS happened to declare — two columns you cannot build a report on. One
selector on the article body makes both measurable, and measured the same way from
one page to the next.Reader Context
loggedIn, subscriptionStatus and goal describe your reader and can only come
from your site:
alkeAnalytics.setDimensions({
loggedIn: true,
subscriptionStatus: 'subscriber',
goal: 'newsletter'
});
subscriptionStatus accepts a closed list of values, so reports stay comparable
across properties:
| Value | Meaning |
|---|---|
anonymous | No account |
registered | Account, no subscription |
trial | Trial period in progress |
subscriber | Active subscription |
expired | Subscription lapsed |
goal records at most one conversion per pageview; setting it twice keeps the last
value. The dashboard’s Goals screens count a conversion once per session, whichever
page carried it, and let you label each value and tune its analysis rules in
Settings › Goals.
When they look at which content leads to a conversion, those screens only consider
the pages read before the page that carried the goal. A goal reached through a
dedicated funnel (offer page, checkout, confirmation) would otherwise explain itself:
give the funnel pages a pageGroup of their own and list it in the goal’s definition
under Settings › Goals, so they are left out of the content analysis while still
counting the conversion.
Loyalty
The collector counts sessions in a cookie and reports the rank of the current visit
as visitSequence. The count is cumulative since the reader’s first visit — it never
decreases — and only resets after 30 days without a visit, since each visit
renews the cookie. It is not a count of visits over the last 30 days: a reader coming
back every three weeks for a year reaches a high rank without having been assiduous
in any given month. newsletter flags readers who arrived from a newsletter (utm_medium=nl,
utm_medium=newsletter, or an email-categorized source) at any point in that same
window. It describes the reader, not the visit: a reader who came through the
newsletter once carries newsletter = yes on every page they read for the next 30
days, whatever brought them back. It is reported as Newsletter reader, and the
question “did this visit come from the newsletter?” is answered by the traffic source
(Email), not by this dimension.
Both are read back synchronously:
alkeAnalytics.getVisitSequence(); // 7
alkeAnalytics.getEngagementLevel(); // 'engaged-regular'
| Visits | Engagement level |
|---|---|
| 0 (unknown) | unknown |
| 1 | fly-by |
| 2–4 | regular-reader |
| 5–9 | engaged-regular |
| 10+ | superfan |
The engagement level is derived from the visit rank in reports, so its thresholds can change without rewriting history.
Both return the same value the report carries, resolved against the same consent
state — not the state read back from the cookie, whose write the CMP may defer. With
storage consent refused, getVisitSequence() returns 0 and getEngagementLevel()
returns unknown. There is no local storage fallback.
The reported dimensions follow the same rule, resolved against the consent state at
send time: visitSequence falls back to 0, so a reader who refused storage lands in
unknown rather than being counted as a first-time fly-by visitor, and newsletter
is simply not declared — the cookie is what carries the answer, so “no consent” is
not “did not come from a newsletter”.
Overriding dimensions
setDimension()
alkeAnalytics.setDimension('level1', 'finance');
| Parameter | Type | Description |
|---|---|---|
name | stringrequired | Dimension name. Also accepts `cd1` to `cd10` as aliases of `setCustomData()`. |
value | string|number|boolean|null|Promiserequired | Value to store. Use `null` to clear. A promise is applied if it resolves before the event is sent, and a promise resolving to `null` clears the dimension just like the synchronous path. Each dimension enforces the shape its column expects, and a value outside it is **rejected** — never truncated or coerced: - `datePublished`, `dateModified` — an integer of **unix seconds**. An ISO string or milliseconds (`Date.now()`) are refused: both would land on the epoch. Passing milliseconds is called out explicitly in the console message. - `loggedIn`, `isAccessibleForFree` — a real boolean. `1` and `0` are refused. - `wordCount` — a positive integer up to 1 000 000, `mediaCount` up to 10 000. These are the exact ceilings the server enforces: above them a value is noise rather than data, and it would land on `0` on arrival. `NaN` and `Infinity` are refused on every numeric dimension. - `subscriptionStatus` — one of the five values listed above. - `countryCode` — exactly two letters. - `coverImage` — an `http`/`https` or relative URL. A relative URL is resolved against the page, since the server requires a scheme, but it is **not** scrubbed: a CDN carries its resizing parameters in the query string. `pageUrl` and `assetUrl` are scrubbed like any other URL the collector reports. - `level1`, `level2`, `level3` (65 characters), `schemaType` and `goal` (64), `abtest` and `pageGroup` (128), `coverImage` (2048) — a string, within that budget. Lengths are counted in code points, so the budget is the same for ASCII, accented and CJK text. - An empty string on `pageUrl`, `assetUrl`, `title`, `pageGroup` or `coverImage` means the same as `null`: the dimension is removed rather than stored blank. On `pageUrl`, `assetUrl` and `title`, the derived value stands again right away — the collector rebuilds them on every event. On `coverImage` it does **not**: the value extracted from the JSON-LD lives in the same store as your override, so clearing it clears the extracted one too, and nothing re-extracts it before the next page boundary. - `abtest` is **trimmed and lowercased on arrival**, like `level1`, `level2` and `level3`: `Paywall-A` and `paywall-a` are one experiment, not two. Pick whichever casing reads best in your code — reports and the A/B test dictionary always show the lowercase form. |
Returns true when the dimension was accepted, false otherwise.
setDimensions()
alkeAnalytics.setDimensions({
level1: 'finance',
level2: 'salaries',
abtest: 'paywall-2026-03-b'
});
setDimensions() is best effort, not all-or-nothing: every valid entry is
applied, and the call returns false when at least one entry was rejected. Check the
console to see which.
Two rejections are worth telling apart there:
- Unknown dimension “…”: it is not settable from the client. — the name is not one of those listed under Available names. The dimensions the server arbitrates fall here too.
- Dimension “…” is derived from the alke_v cookie and cannot be set. —
visitSequenceornewsletter; read them back with the getters instead.
Every other message names the dimension and the shape it expected.
Asynchronous values
A promise is resolved in the background and applied if it lands before the event is
sent, like setLateCustomData(). Set a synchronous default first and refine it
later:
alkeAnalytics.setDimension('subscriptionStatus', 'anonymous');
alkeAnalytics.setDimension('subscriptionStatus', fetchSubscriptionStatus());
The default is kept until the promise resolves. Without a default, a promise that resolves too late simply leaves the dimension out of the event.
Consent-gated values
Dimensions accept the same consent wrapper as custom data:
alkeAnalytics.setDimension('goal',
alkeAnalytics.holdUntilConsent('subscription', 'account', [1, 7])
);
Setting dimensions before setConfig()
Dimensions can be declared before setConfig() runs. setConfig() derives the
automatic dimensions of the first page, but it is not a page boundary: what you
declared beforehand is kept, because the automatic extraction never overwrites an
explicit value.
window.alkeAnalyticsCmd = window.alkeAnalyticsCmd || [];
window.alkeAnalyticsCmd.push(function () {
window.alkeAnalytics.setDimension('subscriptionStatus', 'subscriber');
window.alkeAnalytics.setConfig({ propertyId: 'your-property-uuid' });
});
See Asynchronous loading for why the command queue is the supported way to order these calls.
Declaring a navigation: pushNavigation()
Declare a navigation with pushNavigation(). That call is the page boundary, not
pushPageView():
In a single-page application:
// Page A
alkeAnalytics.setDimensions({ level1: 'finance' });
alkeAnalytics.pushPageView();
// … the reader converts while still on A
alkeAnalytics.setDimension('goal', 'subscription'); // still attributed to A
// The router has swapped the url and rendered page B
alkeAnalytics.pushNavigation(); // A is reported and closed, B is derived
// Page B — the context is its own again
alkeAnalytics.setDimensions({ level1: 'sport' });
alkeAnalytics.pushPageView();
pushNavigation() re-derives the automatic dimensions from the page currently
displayed, so call it once your router has changed the url and rendered. Called too
early it reports the previous page’s hierarchy and metadata, and nothing signals it:
router.afterEach(async () => {
await nextTick(); // url and dom of the new page in place
alkeAnalytics.pushNavigation();
alkeAnalytics.setDimension('goal', null);
alkeAnalytics.pushPageView();
});
Everything declared before the call belongs to the page being left and is
reported with it; everything after belongs to the page to come. You are free to
set a dimension before or after pushPageView() — what matters is which side of
pushNavigation() it falls on.
pushNavigation() also restarts the per-page counters: active time, engagement,
scroll depth and the video sequence go back to zero, so page B never inherits the
engagement of page A.
pushPageView() clears nothing and derives nothing — it opens the report for the
page. That is what lets you set late dimensions, a goal or engagement after the call
and still see them ride that page’s final hit.
A site that loads a real document on each page never needs it — the browser re-initialises the tag by itself.
What the boundary clears
- the content dimensions (hierarchy, article metadata, cover image)
goal- the custom data slots
cd1tocd10 pageUrl,assetUrl,titleandpageGroup- active time, engagement, scroll depth and the video sequence
- the Web Vitals, the navigation timings and the ad creatives already reported
- the page-level video metadata:
setVideoType(),setVideoRebuffer()andsetVideoLoadError()
Nothing survives it, whatever set the value: a dimension you set explicitly for the
page being left does not cross the boundary any more than an extracted one does.
pageUrl, assetUrl and title are re-derived on every event, so clearing them
restores the real url and title: an override you set for the page being left stops
applying at the boundary, it is not carried over. pageGroup has no automatic
value: if you set it, set it again for each page, before its pushPageView().
The values of the page being left are frozen, not merely cleared. A report of
that page still in flight when you call pushNavigation() — held by the consent
queue, or waiting on a bot check — is finalised later with the engagement, active
time and Web Vitals it had at the boundary, and stops counting there. It describes
its own page whenever it happens to leave.
The boundary also opens the next pageviewId: an event emitted between
pushNavigation() and the next pushPageView() already belongs to the page that
follows.
A single-page application that never calls pushNavigation() declares no boundary at
all: the context derived for its first page stands for the whole visit.
What survives, because it describes the reader or the session
loggedIn,subscriptionStatus,newsletter,visitSequenceabtest,countryCode- the traffic source, decided once when the session starts
Automatic extraction re-runs on the next pageview, so the hierarchy and article metadata of the new page are read from the DOM without you doing anything. A value you set explicitly is never overwritten by that extraction.
A promise still resolves onto the page it was set for: if it lands before that page is reported the value is included; if it lands after the navigation, the dimension is simply left out — it is never re-attributed to the following page.
Available names
level1, level2, level3, schemaType, datePublished, dateModified,
isAccessibleForFree, wordCount, mediaCount, coverImage, loggedIn,
subscriptionStatus, goal, abtest, pageUrl, pageGroup, assetUrl, title,
countryCode, and cd1 to cd10.
Names are camelCase, matching the rest of the SDK. They are the same identifiers the collector puts on the wire, so there is no second vocabulary to learn.
countryCode is accepted here and takes precedence over the CDN geolocation, exactly
like the countryCode configuration option it mirrors.
pageUrl, assetUrl and title are settable too, and your value wins: the
collector derives them when it builds an event, but the dimensions you declared are
overlaid on top just before the report is sent. They are page-scoped, so set them
again after each pushNavigation(). assetUrl carries no value on a pageview unless
you set one.
videoplay report carries its own assetUrl (the media URL) and title (the
video title). The overlay applies to it like to any other event, so an assetUrl or
a title you set for the page replaces the video’s own value on that report.
Leave the two unset unless you mean to override them everywhere, video included.Read-only dimensions
visitSequence and newsletter are reported, and reportable on, but cannot be set:
setDimension() rejects them and returns false. Both are derived from the alke_v
cookie, and you read them back with getVisitSequence() and getEngagementLevel() —
a value you pushed would disagree with what those getters answer for the same reader.
Dimensions the server arbitrates — bot detection, device and browser families — are not settable either. They are computed by the collector and sent, but the server normalizes them against its own allowlists, so overriding them would change nothing.