Ad Server Data Transfer
Deliver Google Ad Manager Data Transfer files to Alke Analytics so ad revenue can be attributed to individual pageviews.
Overview
Alke Analytics joins your ad server events to the pageviews collected by the SDK, so that revenue, impressions, requests and viewability become filterable and pivotable on every analytics dimension (page, country, device, traffic source, custom dimension…).
This page is the contract for that delivery. It covers the storage location, the directory layout, the file naming, and the exact column names and types expected for each file type.
1. Prerequisite — the alke_id custom targeting key
Google Ad Manager only writes custom targeting keys into Data Transfer files if those keys are declared in the network. Before any delivery is useful:
- Create a custom targeting key named exactly
alke_idin Google Ad Manager (Inventory → Key-values), of type free-form text. - Make sure your ad request tags forward the value produced by the Alke collector.
- Confirm that the key appears in the
customtargetingcolumn of a fresh file.
The value is a lowercase UUID, 36 characters:
alke_id=536c8820-94a6-11f1-836e-318a0522cae2
Alke Analytics extracts it with the pattern alke_id=([0-9a-f\-]{36}) and discards
every row that does not carry it. Uppercase hexadecimal, a different key name, or a
truncated value all result in silent zero-revenue ingestion.
2. Storage and access
| Protocol | S3-compatible object storage (s3://) or Google Cloud Storage (gs://) |
| Bucket and root prefix | Provided by Alke Tech |
| Credentials | Provided by Alke Tech; read access is configured on the Alke side |
| Allowed characters in bucket and prefix | A-Z a-z 0-9 . _ - / = |
Alke Analytics reads the files directly from that storage. Nothing is pushed to an API, and no other transfer channel is supported.
3. Directory layout
Every file must be reachable at exactly this path, relative to the agreed root prefix:
<file-type>/year=YYYY/month=MM/day=DD/<file-name>.parquet
Example:
networkbackfillimpressions/year=2026/month=08/day=12/part-00007.parquet
Rules:
<file-type>is one of the eight values in section 4, lowercase, no separators.year,monthanddayare zero-padded (month=08, notmonth=8).- The date in the path is the event date — the day the impression, click or request occurred. It is not the date the export ran. See section 7.
- There is no additional nesting. The
.parquetfiles sit directly in theday=DDdirectory. - Any number of files per day is fine; they are all read.
File names
File names must match ^[A-Za-z0-9._-]+\.parquet$ — letters, digits, dot, underscore
and hyphen only. Any other character (space, quote, %, +, …) causes the file to be
rejected.
File immutability
Alke Analytics tracks ingestion per file path and contributes each file’s revenue exactly once. A file that has already been ingested must never be modified in place.
- To add data, deliver a new file in the same directory.
- To correct data already delivered, contact Alke Tech so the affected day can be rebuilt. Overwriting a file silently produces either missing or double-counted revenue.
4. File types
Deliver the file types your setup produces. Direct and backfill (Ad Exchange / Open Bidding) inventory are separate types and must not be merged.
| Directory name | Content |
|---|---|
networkimpressions | Direct-sold impressions |
networkbackfillimpressions | Backfill (Ad Exchange, Open Bidding, EBDA) impressions |
networkclicks | Direct-sold clicks |
networkbackfillclicks | Backfill clicks |
networkrequests | Direct ad requests |
networkbackfillrequests | Backfill ad requests |
networkactiveviews | Direct viewability measurements |
networkbackfillactiveviews | Backfill viewability measurements |
Missing a type degrades the corresponding metric only: without networkrequests there
is no fill-rate, without networkactiveviews there is no viewability rate. Impressions
are required for any revenue at all.
5. File format
| Format | Apache Parquet |
| Compression | snappy, lz4, brotli, zstd, gzip, or uncompressed (none) |
| Column names | lowercase, exactly as listed in section 6 |
| Target file size | 128–512 MB per file |
A column named Time or TIME is not the same as time and will not be found.
Systems that uppercase identifiers by default — Snowflake, Oracle, Hive — must quote
the aliases to preserve lowercase.
A Parquet file whose columns are named _COL_0, _COL_1, … carries no usable schema
and cannot be ingested. This happens when an export tool unloads an unaliased
SELECT *; alias every column explicitly.
6. Column specifications
Columns not listed here are ignored — extra columns are harmless. Every listed column
must be physically present in the file: the schema is inferred from the Parquet file
and each column is referenced unconditionally, so a missing column fails the ingestion of
the whole file type (UNKNOWN_IDENTIFIER). “Optional” below never means the column may be
absent — it means only its value may be empty ('' or 0); the column itself must
still be emitted.
An empty value must be an empty string '' or 0, never NULL. A NULL is not
rejected: every string and id column is read into a non-nullable field that carries a
default, and ClickHouse’s default behaviour writes a NULL as that default ('' / 0)
rather than raising an error. This is worse than a missing column — there is no
UNKNOWN_IDENTIFIER to notice — because it is silent: a NULL revenue is recorded as 0
(lost revenue, not a failed import). Use COALESCE(col, '') / COALESCE(col, 0) in your
export for any column that can be null at source.
Common types
- timestamp string — the event time as text,
YYYY-MM-DDTHH:MM:SS. The variantYYYY-MM-DD-HH:MM:SSis also accepted. Must be a string, not a ParquetTIMESTAMP; the first 10 characters are read as the date and characters 12 onward as the time. - id — integer, or a string of digits. Both are accepted. Decimal types must have scale 0.
- targeting string — semicolon-separated
key=valuepairs, e.g.device=smartphone;alke_id=536c8820-94a6-11f1-836e-318a0522cae2;pagetype=article.
networkimpressions
| Column | Type | Notes |
|---|---|---|
time | timestamp string | |
impressionid | string | Column required; value optional — emit '' if your source has no equivalent |
customtargeting | targeting string | Must contain alke_id= |
refererurl | string | Full page URL |
adunitid | id | |
orderid | id | |
lineitemid | id | Used to price the impression from the line item |
creativeid | id | |
creativesize | string | e.g. 300x250 |
advertiserid | id | |
product | string | e.g. Ad Server |
country | string | Country name as produced by GAM, e.g. France |
devicecategory | string | e.g. Smartphone |
activevieweligiblecount | integer | 0 or 1 |
activeviewmeasurablecount | integer | 0 or 1 |
Direct impressions carry no revenue column: they are valued from the line item’s cost type and cost per unit, which Alke Analytics reads from the Ad Manager API.
networkbackfillimpressions
All columns of networkimpressions, except creativesize which is replaced by
creativesizedelivered, plus:
| Column | Type | Notes |
|---|---|---|
creativesizedelivered | string | Size actually delivered |
estimatedbackfillrevenue | floating point | Per impression, in the network currency — not a CPM, not micros. A value of 0.0002 means 0.0002 € for that single impression. |
advertiser | string | Advertiser name |
buyer | string | Buyer / DSP name |
yieldgroupnames | string | Pipe-separated when multiple |
dealid | id | Column required; value optional — emit 0 when absent |
dealtype | string | Column required; value optional — emit '' when absent |
The currency is configured per publisher on the Alke side and is not read from the file. Confirm it with Alke Tech during onboarding.
networkclicks and networkbackfillclicks
| Column | Type | Notes |
|---|---|---|
time | timestamp string | |
customtargeting | targeting string | Must contain alke_id= |
lineitemid | id | Used to price CPC line items |
country | string | |
devicecategory | string |
networkrequests and networkbackfillrequests
| Column | Type | Notes |
|---|---|---|
time | timestamp string | |
customtargeting | targeting string | Must contain alke_id= |
adunitid | id | |
isfilledrequest | boolean | true when the request was filled |
networkactiveviews and networkbackfillactiveviews
| Column | Type | Notes |
|---|---|---|
time | timestamp string | |
customtargeting | targeting string | Must contain alke_id= |
adunitid | id | |
lineitemid | id | |
activeviewviewablecount | integer | 0 or 1 |
activeviewmeasurablecount | integer | 0 or 1 |
7. Delivery cadence and the attribution window
Alke Analytics attaches an ad event to a pageview day within a window of plus or minus one day around the date in the file path. This absorbs the normal skew between the client clock and the ad server clock.
Two consequences:
- Partition by event date, not by export date. If files are laid out by the time the export ran, every event is shifted by the export lag, and the whole feed rides the edge of the window. A one-day lag is absorbed; a two-day lag loses revenue silently.
- A late file is fine, a misfiled file is not. Delivering the
day=12directory on the 14th works perfectly — the path is what matters, not the delivery time.
There is no required cadence. Delivering continuously as events land is just as good as
one batch per day — files accumulate in the day=DD directory of their event date and
are picked up as they arrive.
Source retention
Keep delivered files available for at least 30 days. Alke Analytics retains the raw ad server facts for 7 days internally; rebuilding a day older than that requires re-reading the source files.
8. Validation checklist
Before the first production delivery, confirm on one day of data:
- The path resolves to
<file-type>/year=YYYY/month=MM/day=DD/*.parquetunder the agreed root prefix. - File names match
^[A-Za-z0-9._-]+\.parquet$. - Column names are lowercase and match section 6 exactly — no
_COL_0, noTIME. -
timeis a string, and its first 10 characters are the date. -
customtargetingis populated, and the share of rows containingalke_id=matches the share of traffic where the Alke collector is deployed. - The extracted
alke_idvalues are 36-character lowercase UUIDs. - The event dates inside a
day=DDdirectory are that day, not the export day. -
estimatedbackfillrevenuemagnitudes are per impression (typically 10⁻⁵ to 10⁻¹).
Alke Tech runs the same checks on the first delivery and reports back before the feed is switched on.