Content Dimensions

Report on content hierarchy, article metadata and reader context as first class dimensions.

Overview

Availability

CapabilityWebiOS / tvOSAndroid / Android TVTizen
setDimension() / setDimensions()✅✅✅✅
getVisitSequence() / getEngagementLevel()✅✅✅✅
newsletter from the opening link (handleOpenURL())✅✅✅— (left undeclared)
Automatic extraction (hierarchy, JSON-LD, word and media counts)✅———
pushNavigation()✅✅✅✅

The automatic extraction is web only. It reads the URL path, the page’s JSON-LD and the DOM, none of which exists in a native application. On the native SDKs every content dimension is pushed explicitly by your app, and a dimension you do not push simply keeps its default value.

pageUrl and title are not dimensions on the native SDKs: they are already the url: and title: arguments of pushPageView. assetUrl is settable there, because nothing derives it for you.

Beyond the 10 custom data slots, Alke Analytics ships a set of named dimensions that model editorial content and reader context directly. Most of them are filled in automatically by the web collector; every one of them can be overridden by your site.

DimensionFilled byDescription
level1, level2, level3AutomaticContent hierarchy derived from the URL path
schemaTypeAutomaticSchema.org @type of the page
datePublished, dateModifiedAutomaticPublication and update dates
isAccessibleForFreeAutomaticWhether the article is free to read — three states: declared free, declared paywalled, or not declared
wordCount, mediaCountAutomaticSize of the article body
coverImageAutomaticCover image URL
loggedInYour siteWhether the reader is authenticated
subscriptionStatusYour siteSubscription state of the reader
goalYour siteConversion reached on the page
abtestYour siteIdentifier of the running experiment
pageGroupYour siteFree-form grouping of the page — no automatic value
newsletter, visitSequenceAutomaticReader loyalty, from the visit cookie. Not settable — read them back with the getters

Content Hierarchy

The collector splits the URL path on / and drops the empty segments. The last segment is treated as the article slug and dropped only when the page is a piece of content — that is, when the schema.org @type of the selected node (see Article Metadata) is a content type. With no structured data, or with a generic page wrapper such as WebPage or CollectionPage, no segment is dropped and the path is taken as it is. A trailing slash plays no part: it is the default WordPress permalink, so keying on it dropped the section on an article URL and pushed the slug into level2/level3 everywhere else.

The remaining segments feed level1, level2 and level3.

Path@type of the selected nodelevel1level2level3
/economy/salaries/my-article-84f46e03NewsArticleeconomysalaries—
/economy/salaries/my-article-84f46e03none, or WebPageeconomysalariesmy-article-84f46e03
/economy/salaries/CollectionPageeconomysalaries—
/a/b/c/dArticleabc
/a/b/c/dnoneabc
/my-articleArticle———
/my-articlenonemy-article——

On a page carrying no structured data, level3 can therefore hold a slug. That is accepted — the column is sized for it.

Segments are decoded, lowercased, stripped of control characters and of any / or \ a decoded %2F would put back, then cut to 65 characters. A harmonized level you push yourself goes through the same normalization server-side, so it groups with a derived one.

If your URLs do not map cleanly onto your editorial taxonomy, push a harmonized version instead — see Overriding dimensions.

Page grouping

pageGroup is a free-form grouping of your own choosing, orthogonal to the three hierarchy levels: a template name, a franchise, an editorial format. It has no automatic value — nothing is reported unless you set it:

alkeAnalytics.setDimension('pageGroup', 'longform-template');

Up to 128 characters; an empty string removes it, like null. It describes one page, so it is cleared at every page boundary — see what the boundary clears.

Article Metadata

Article metadata is read from the page’s JSON-LD blocks (<script type="application/ld+json">), CDATA-wrapped ones included. The graph is walked depth-first through @graph, mainEntity, mainEntityOfPage, hasPart and itemListElement, up to 6 levels deep — WebPage → mainEntity → NewsArticle is the nominal shape outside SEO plugins, and stopping at the wrapper would report a page with no metadata at all.

“That page” is your declared <link rel="canonical"> when there is one, and the current URL without its query string and fragment otherwise — a canonical URL carries neither, so comparing against the raw address would fail the identity test below on every campaign arrival (?utm_source=…, ?fbclid=…).

The node describing the page is then picked in this order:

  1. The node that claims to be this page — an @id, url or mainEntityOfPage equal to that URL (fragment and trailing slash ignored), preferring a content type over a page wrapper. Without this, a section page listing articles, or an article followed by a block of recommendations, would report the first listed item’s date and paywall as its own.
  2. Otherwise, the single candidate of its family: the content types first, then the generic page wrappers. Both families are closed lists, given below. Several candidates, none of which claims the URL, means the page lists content rather than being it, so the next family is tried.
  3. Failing both, nothing is reported: an Organization or a WebSite node describes your site, not the page in front of the reader. There is no fallback onto “the first typed node”.

That ordering matters: SEO plugins emit a WebPage node alongside the article one, in an order you do not control. Preferring the content node means the same page extracts the same data whichever way the plugin writes its @graph.

Membership is decided by exact name, not by schema.org inheritance: a type absent from both lists below is never selected, however specific it is. The same content list decides whether the last URL segment is dropped as an article slug — see Content Hierarchy.

Content types (30) — the page is the content

Article, NewsArticle, BlogPosting, LiveBlogPosting, TechArticle, ScholarlyArticle, SocialMediaPosting, AdvertiserContentArticle, AnalysisNewsArticle, AskPublicNewsArticle, BackgroundNewsArticle, OpinionNewsArticle, ReportageNewsArticle, ReviewNewsArticle, SatiricalArticle, Report, Review, Recipe, HowTo, Course, VideoObject, AudioObject, PodcastEpisode, Episode, Movie, Book, RealEstateListing, MediaGallery, ImageGallery, VideoGallery

Page wrappers (11) — the page describes or lists content

WebPage, ItemPage, CollectionPage, ProfilePage, AboutPage, ContactPage, CheckoutPage, FAQPage, QAPage, SearchResultsPage, CreativeWork

The schema.org namespace is recognised in its http/https and www. forms, so a canonical "@type": "https://schema.org/NewsArticle" is read like the compact one. A schema: CURIE is not resolved — its binding lives in @context.

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "NewsArticle",
  "datePublished": "2026-03-04T08:30:00+01:00",
  "dateModified": "2026-03-04T10:00:00Z",
  "isAccessibleForFree": false,
  "wordCount": 812,
  "image": "https://cdn.example.org/cover.jpg"
}
</script>

Nothing found means nothing reported: the dimension simply keeps its default value.

Three states, not two

isAccessibleForFree, loggedIn and newsletter are stored as three states — declared true, declared false, and not declared — rather than as booleans.

isAccessibleForFree is only reported when the page actually declares it. A page whose paywall could not be read reports nothing and is no longer counted as paywalled, which is what a boolean whose default was “paid” used to do to an entire site with no paywall. On loggedIn and newsletter, the same distinction separates “the site declared false” from “the site does not implement this dimension”.

These forms are accepted for isAccessibleForFree: a JSON boolean, "True"/"False" in any case, the enumeration IRIs http(s)://schema.org/True|False, the integers 0/1 (and their string forms), a one-element array, and an @value wrapper around any of these. As a last resort, a hasPart[] declaring isAccessibleForFree: false — Google’s documented pattern for a partial paywall — counts as a false declaration for the page. Anything else counts as not declared.

Dates and word count

datePublished and dateModified are read as ISO 8601 only (YYYY-MM-DD, with or without a time component). Anything else is ignored rather than guessed: 04/03/2026 would otherwise parse into a valid and wrong date. With no offset declared, the value is read as UTC, so the same article does not carry a different date per reader.

The JSON-LD wordCount is only read when it is a plain integer: "1,234" is refused, because a lenient parse stops at the separator and reports 1 — a plausible integer nothing downstream can catch.

image may be a URL, an ImageObject, an array of either, or an @id reference to a node declared elsewhere in the graph. A reference is only dereferenced when the target really is an image, so a mis-wired image pointing at the WebPage node does not store the page URL as a cover. thumbnailUrl is used as a fallback.

Counting words and media

By default the word count comes from the JSON-LD wordCount, falling back to counting articleBody. Media are not counted, because nothing indicates where the article body is in the DOM.

Two consequences worth knowing before you build a report on these two columns:

  • mediaCount stays at 0 for every page until you configure contentSelector. Nothing distinguishes “this article has no media” from “media were never measured”, so treat the column as empty on a property that does not set the selector.
  • wordCount can come from three different instruments — the DOM, the JSON-LD wordCount declared by the CMS, or a count of articleBody — which do not share a definition of a word. Comparing averages across properties compares instruments; within one property, configure contentSelector so every page is measured the same way.

Configure contentSelector to measure the DOM instead — words from the node’s text with script, style, noscript and template removed, so an embed or an inline JSON-LD sitting in the article body is not counted as prose; media from the img, picture, video, audio and iframe elements it contains:

alkeAnalytics.setConfig({
  propertyId: 'your-property-id',
  contentSelector: 'article.body'
});

Three kinds of element are left out of the media count on purpose: an img nested in a picture (same media as its parent), an img/video/audio with no source at all (a lazy-load placeholder), and a cross-origin iframe — in an article body that is more often an ad slot, a comment widget or a signup form than a media.

If the selector matches nothing, the collector falls back to the JSON-LD as above and logs a warning.

Set a content selector on every editorial property
Without it, mediaCount stays at 0 on every page and wordCount comes from whatever the CMS happened to declare — two columns you cannot build a report on. One selector on the article body makes both measurable, and measured the same way from one page to the next.

Reader Context

loggedIn, subscriptionStatus and goal describe your reader and can only come from your site:

alkeAnalytics.setDimensions({
  loggedIn: true,
  subscriptionStatus: 'subscriber',
  goal: 'newsletter'
});

subscriptionStatus accepts a closed list of values, so reports stay comparable across properties:

ValueMeaning
anonymousNo account
registeredAccount, no subscription
trialTrial period in progress
subscriberActive subscription
expiredSubscription lapsed

goal records at most one conversion per pageview; setting it twice keeps the last value. The dashboard’s Goals screens count a conversion once per session, whichever page carried it, and let you label each value and tune its analysis rules in Settings › Goals.

When they look at which content leads to a conversion, those screens only consider the pages read before the page that carried the goal. A goal reached through a dedicated funnel (offer page, checkout, confirmation) would otherwise explain itself: give the funnel pages a pageGroup of their own and list it in the goal’s definition under Settings › Goals, so they are left out of the content analysis while still counting the conversion.

Loyalty

The collector counts sessions in a cookie and reports the rank of the current visit as visitSequence. The count is cumulative since the reader’s first visit — it never decreases — and only resets after 30 days without a visit, since each visit renews the cookie. It is not a count of visits over the last 30 days: a reader coming back every three weeks for a year reaches a high rank without having been assiduous in any given month. newsletter flags readers who arrived from a newsletter (utm_medium=nl, utm_medium=newsletter, or an email-categorized source) at any point in that same window. It describes the reader, not the visit: a reader who came through the newsletter once carries newsletter = yes on every page they read for the next 30 days, whatever brought them back. It is reported as Newsletter reader, and the question “did this visit come from the newsletter?” is answered by the traffic source (Email), not by this dimension.

Both are read back synchronously:

alkeAnalytics.getVisitSequence();   // 7
alkeAnalytics.getEngagementLevel(); // 'engaged-regular'
VisitsEngagement level
0 (unknown)unknown
1fly-by
2–4regular-reader
5–9engaged-regular
10+superfan

The engagement level is derived from the visit rank in reports, so its thresholds can change without rewriting history.

Both return the same value the report carries, resolved against the same consent state — not the state read back from the cookie, whose write the CMP may defer. With storage consent refused, getVisitSequence() returns 0 and getEngagementLevel() returns unknown. There is no local storage fallback.

The reported dimensions follow the same rule, resolved against the consent state at send time: visitSequence falls back to 0, so a reader who refused storage lands in unknown rather than being counted as a first-time fly-by visitor, and newsletter is simply not declared — the cookie is what carries the answer, so “no consent” is not “did not come from a newsletter”.

Overriding dimensions

setDimension()

alkeAnalytics.setDimension('level1', 'finance');
ParameterTypeDescription
namestringrequiredDimension name. Also accepts `cd1` to `cd10` as aliases of `setCustomData()`.
valuestring|number|boolean|null|PromiserequiredValue to store. Use `null` to clear. A promise is applied if it resolves before the event is sent, and a promise resolving to `null` clears the dimension just like the synchronous path. Each dimension enforces the shape its column expects, and a value outside it is **rejected** — never truncated or coerced: - `datePublished`, `dateModified` — an integer of **unix seconds**. An ISO string or milliseconds (`Date.now()`) are refused: both would land on the epoch. Passing milliseconds is called out explicitly in the console message. - `loggedIn`, `isAccessibleForFree` — a real boolean. `1` and `0` are refused. - `wordCount` — a positive integer up to 1 000 000, `mediaCount` up to 10 000. These are the exact ceilings the server enforces: above them a value is noise rather than data, and it would land on `0` on arrival. `NaN` and `Infinity` are refused on every numeric dimension. - `subscriptionStatus` — one of the five values listed above. - `countryCode` — exactly two letters. - `coverImage` — an `http`/`https` or relative URL. A relative URL is resolved against the page, since the server requires a scheme, but it is **not** scrubbed: a CDN carries its resizing parameters in the query string. `pageUrl` and `assetUrl` are scrubbed like any other URL the collector reports. - `level1`, `level2`, `level3` (65 characters), `schemaType` and `goal` (64), `abtest` and `pageGroup` (128), `coverImage` (2048) — a string, within that budget. Lengths are counted in code points, so the budget is the same for ASCII, accented and CJK text. - An empty string on `pageUrl`, `assetUrl`, `title`, `pageGroup` or `coverImage` means the same as `null`: the dimension is removed rather than stored blank. On `pageUrl`, `assetUrl` and `title`, the derived value stands again right away — the collector rebuilds them on every event. On `coverImage` it does **not**: the value extracted from the JSON-LD lives in the same store as your override, so clearing it clears the extracted one too, and nothing re-extracts it before the next page boundary. - `abtest` is **trimmed and lowercased on arrival**, like `level1`, `level2` and `level3`: `Paywall-A` and `paywall-a` are one experiment, not two. Pick whichever casing reads best in your code — reports and the A/B test dictionary always show the lowercase form.

Returns true when the dimension was accepted, false otherwise.

setDimensions()

alkeAnalytics.setDimensions({
  level1: 'finance',
  level2: 'salaries',
  abtest: 'paywall-2026-03-b'
});

setDimensions() is best effort, not all-or-nothing: every valid entry is applied, and the call returns false when at least one entry was rejected. Check the console to see which.

Two rejections are worth telling apart there:

  • Unknown dimension “…”: it is not settable from the client. — the name is not one of those listed under Available names. The dimensions the server arbitrates fall here too.
  • Dimension “…” is derived from the alke_v cookie and cannot be set. — visitSequence or newsletter; read them back with the getters instead.

Every other message names the dimension and the shape it expected.

Asynchronous values

A promise is resolved in the background and applied if it lands before the event is sent, like setLateCustomData(). Set a synchronous default first and refine it later:

alkeAnalytics.setDimension('subscriptionStatus', 'anonymous');
alkeAnalytics.setDimension('subscriptionStatus', fetchSubscriptionStatus());

The default is kept until the promise resolves. Without a default, a promise that resolves too late simply leaves the dimension out of the event.

Dimensions accept the same consent wrapper as custom data:

alkeAnalytics.setDimension('goal',
  alkeAnalytics.holdUntilConsent('subscription', 'account', [1, 7])
);

Setting dimensions before setConfig()

Dimensions can be declared before setConfig() runs. setConfig() derives the automatic dimensions of the first page, but it is not a page boundary: what you declared beforehand is kept, because the automatic extraction never overwrites an explicit value.

window.alkeAnalyticsCmd = window.alkeAnalyticsCmd || [];
window.alkeAnalyticsCmd.push(function () {
    window.alkeAnalytics.setDimension('subscriptionStatus', 'subscriber');
    window.alkeAnalytics.setConfig({ propertyId: 'your-property-uuid' });
});

See Asynchronous loading for why the command queue is the supported way to order these calls.

Declaring a navigation: pushNavigation()

Declare a navigation with pushNavigation(). That call is the page boundary, not pushPageView():

In a single-page application:

// Page A
alkeAnalytics.setDimensions({ level1: 'finance' });
alkeAnalytics.pushPageView();

// … the reader converts while still on A
alkeAnalytics.setDimension('goal', 'subscription');   // still attributed to A

// The router has swapped the url and rendered page B
alkeAnalytics.pushNavigation();          // A is reported and closed, B is derived

// Page B — the context is its own again
alkeAnalytics.setDimensions({ level1: 'sport' });
alkeAnalytics.pushPageView();

pushNavigation() re-derives the automatic dimensions from the page currently displayed, so call it once your router has changed the url and rendered. Called too early it reports the previous page’s hierarchy and metadata, and nothing signals it:

router.afterEach(async () => {
    await nextTick();                // url and dom of the new page in place
    alkeAnalytics.pushNavigation();
    alkeAnalytics.setDimension('goal', null);
    alkeAnalytics.pushPageView();
});

Everything declared before the call belongs to the page being left and is reported with it; everything after belongs to the page to come. You are free to set a dimension before or after pushPageView() — what matters is which side of pushNavigation() it falls on.

pushNavigation() also restarts the per-page counters: active time, engagement, scroll depth and the video sequence go back to zero, so page B never inherits the engagement of page A.

pushPageView() clears nothing and derives nothing — it opens the report for the page. That is what lets you set late dimensions, a goal or engagement after the call and still see them ride that page’s final hit.

A site that loads a real document on each page never needs it — the browser re-initialises the tag by itself.

What the boundary clears

  • the content dimensions (hierarchy, article metadata, cover image)
  • goal
  • the custom data slots cd1 to cd10
  • pageUrl, assetUrl, title and pageGroup
  • active time, engagement, scroll depth and the video sequence
  • the Web Vitals, the navigation timings and the ad creatives already reported
  • the page-level video metadata: setVideoType(), setVideoRebuffer() and setVideoLoadError()

Nothing survives it, whatever set the value: a dimension you set explicitly for the page being left does not cross the boundary any more than an extracted one does. pageUrl, assetUrl and title are re-derived on every event, so clearing them restores the real url and title: an override you set for the page being left stops applying at the boundary, it is not carried over. pageGroup has no automatic value: if you set it, set it again for each page, before its pushPageView().

The values of the page being left are frozen, not merely cleared. A report of that page still in flight when you call pushNavigation() — held by the consent queue, or waiting on a bot check — is finalised later with the engagement, active time and Web Vitals it had at the boundary, and stops counting there. It describes its own page whenever it happens to leave.

The boundary also opens the next pageviewId: an event emitted between pushNavigation() and the next pushPageView() already belongs to the page that follows.

A single-page application that never calls pushNavigation() declares no boundary at all: the context derived for its first page stands for the whole visit.

What survives, because it describes the reader or the session

  • loggedIn, subscriptionStatus, newsletter, visitSequence
  • abtest, countryCode
  • the traffic source, decided once when the session starts

Automatic extraction re-runs on the next pageview, so the hierarchy and article metadata of the new page are read from the DOM without you doing anything. A value you set explicitly is never overwritten by that extraction.

A promise still resolves onto the page it was set for: if it lands before that page is reported the value is included; if it lands after the navigation, the dimension is simply left out — it is never re-attributed to the following page.

Available names

level1, level2, level3, schemaType, datePublished, dateModified, isAccessibleForFree, wordCount, mediaCount, coverImage, loggedIn, subscriptionStatus, goal, abtest, pageUrl, pageGroup, assetUrl, title, countryCode, and cd1 to cd10.

Names are camelCase, matching the rest of the SDK. They are the same identifiers the collector puts on the wire, so there is no second vocabulary to learn.

countryCode is accepted here and takes precedence over the CDN geolocation, exactly like the countryCode configuration option it mirrors.

pageUrl, assetUrl and title are settable too, and your value wins: the collector derives them when it builds an event, but the dimensions you declared are overlaid on top just before the report is sent. They are page-scoped, so set them again after each pushNavigation(). assetUrl carries no value on a pageview unless you set one.

These three reach video events as well
A videoplay report carries its own assetUrl (the media URL) and title (the video title). The overlay applies to it like to any other event, so an assetUrl or a title you set for the page replaces the video’s own value on that report. Leave the two unset unless you mean to override them everywhere, video included.

Read-only dimensions

visitSequence and newsletter are reported, and reportable on, but cannot be set: setDimension() rejects them and returns false. Both are derived from the alke_v cookie, and you read them back with getVisitSequence() and getEngagementLevel() — a value you pushed would disagree with what those getters answer for the same reader.

Dimensions the server arbitrates — bot detection, device and browser families — are not settable either. They are computed by the collector and sent, but the server normalizes them against its own allowlists, so overriding them would change nothing.