Cluster
Un cluster è un elenco salvato di siti web. Con l'API puoi crearne uno da una ricerca o da un tuo elenco, combinare più cluster e ricavare valori dal codice sorgente delle pagine di ogni sito che contiene: contatti, profili social, ID di analytics e di tag, tutto ciò che un'espressione regolare può trovare.
Gli endpoint
| Endpoint | Metodo | Cosa fa |
|---|---|---|
/v1/clusters | GET | I cluster dell'account, dal più recente: ID, nome, numero di domini, data di creazione. |
/v1/clusters | POST | Crea un cluster da una ricerca (query) o da un elenco (domains); name è facoltativo. |
/v1/clusters/{id} | GET | Il cluster e una pagina dei suoi domini: page, per_page (fino a 10.000), format=txt per avere un dominio per riga. |
/v1/clusters/combine | POST | Un nuovo cluster da cluster esistenti: operation and, or o diff, clusters, cioè un elenco di ID. |
/v1/clusters/{id}/extract | GET, POST | Estrae valori con presets e regex, una parte alla volta: offset, limit (fino a 1.000 siti per chiamata), format json, xml o csv. |
/v1/clusters/presets | GET | Le espressioni già pronte, con i loro pattern. |
/v1/clusters/{id}/rename | POST | Nuovo name. |
/v1/clusters/{id}/delete | POST | Elimina il cluster definitivamente. |
Tutto ciò che modifica un cluster è una richiesta POST con un corpo JSON, senza
PUT né DELETE: così qualsiasi client in grado di cercare può anche gestire i
cluster. Il token, il limite di frequenza e il formato degli errori sono gli
stessi della ricerca; i numeri dei cluster
nella risposta di /v1/account mostrano quanti cluster hai e quanti
punti di estrazione ti restano oggi.
Creare un cluster
Da una ricerca: i risultati della query, fino al numero massimo di righe
per ricerca del tuo piano, diventano il cluster. Costa una ricerca, come
/v1/search.
curl https://api.publicwww.com/v1/clusters \
-H "Authorization: Bearer $PUBLICWWW_KEY" \
-H "Content-Type: application/json" \
-d '{"query": "\"googletagmanager.com/gtm.js\"", "name": "GTM sites"}'
{"id": 7, "name": "GTM sites", "size": 100000, "created": "2026-10-10T20:21:30Z",
"query": "\"googletagmanager.com/gtm.js\"", "total": 2412577, "index_complete": true}
Da un elenco: domini o URL, come array JSON o uno per riga. Si
conservano solo i siti presenti nell'indice: submitted indica
quanti ne hai inviati, size quanti sono entrati nel cluster. Un
elenco non costa nulla.
curl https://api.publicwww.com/v1/clusters \
-H "Authorization: Bearer $PUBLICWWW_KEY" \
-H "Content-Type: application/json" \
-d '{"domains": ["example.com", "https://www.example.org/about"], "name": "Prospects"}'
Il piano limita il numero di domini in un cluster, e un account conserva fino a
100 cluster. Se ce ne sono già 100, la creazione di un altro risponde
409 cluster_limit: non eliminiamo nulla al posto tuo, quindi
elimina prima quelli che non ti servono più.
Combinare i cluster
curl https://api.publicwww.com/v1/clusters/combine \
-H "Authorization: Bearer $PUBLICWWW_KEY" \
-H "Content-Type: application/json" \
-d '{"operation": "diff", "clusters": [7, 3], "name": "GTM, not yet contacted"}'
and tiene i domini presenti in tutti i cluster, or
quelli presenti in almeno uno e diff quelli del primo cluster che
non sono nel secondo (in questo caso i cluster sono esattamente due). Combinare
non costa nulla.
Estrarre i dati
L'estrazione legge le pagine indicizzate di ogni sito del cluster e restituisce ciò che le espressioni catturano, una colonna per espressione. Il modo più semplice è un modello predefinito:
| Modello predefinito | Cosa ricava |
|---|---|
email | Indirizzi dai link mailto: |
phone | Numeri dai link tel: |
whatsapp, telegram, skype | Numeri WhatsApp, nomi utente Telegram e nomi Skype, dai rispettivi link |
facebook, instagram, twitter, linkedin | Link ai profili social del sito (twitter accetta anche x.com, linkedin le pagine aziendali) |
gtm, ga4, ua | ID del container di Google Tag Manager, di Google Analytics 4 e di Universal Analytics |
hotjar | ID del sito Hotjar |
adsense | ID publisher AdSense |
bitcoin | Indirizzi dai link di pagamento bitcoin: |
Oppure scrivi un'espressione tua: un'espressione regolare tra barre (o barre
verticali) con i flag facoltativi i, m, s,
u, lunga al massimo 200 caratteri. Il primo gruppo di cattura è il
valore, con la stessa regola di
snipexp:. Fino a dieci
espressioni per chiamata, modelli predefiniti compresi.
curl https://api.publicwww.com/v1/clusters/7/extract \
-H "Authorization: Bearer $PUBLICWWW_KEY" \
-H "Content-Type: application/json" \
-d '{"presets": ["gtm", "email"], "regex": ["/data-site-id=\"([0-9]+)\"/i"], "limit": 1000}'
{
"cluster": 7, "name": "GTM sites", "size": 100000,
"offset": 0, "scanned": 1000, "in_index": 1000, "with_matches": 941,
"next_offset": 1000,
"regex": ["/(GTM-[A-Z0-9]{4,10})\\b/", "/mailto:(...)/i", "/data-site-id=\"([0-9]+)\"/i"],
"points_used": 1834.2, "points_left": 98165,
"rows": [
{ "domain": "example.com", "values": [["GTM-AB12CD"], ["info@example.com"], []], "matches": 2 }
]
}
Una parte alla volta. Una chiamata scorre fino a 1.000 siti del cluster, a
partire da offset. Richiama con offset impostato a
next_offset finché non torna null. I siti senza
corrispondenze vengono omessi; con skip_empty=0 compaiono anche loro.
format=csv restituisce una riga per sito (il dominio, poi una colonna
per espressione), con l'offset successivo nell'header X-Next-Offset.
Punti. L'estrazione consuma la dotazione giornaliera di estrazione del tuo
piano: ogni sito trovato nell'indice costa tanti punti quanti sono i valori
trovati, 0,1 se non ce ne sono. Quando i punti finiscono, la chiamata si ferma
prima con "stopped": "extract_quota_exceeded" e restituisce ciò che
ha raccolto fin lì; una chiamata quando non resta più nulla riceve
429 extract_quota_exceeded. Prova prima un'espressione con un
limit piccolo: vedi i limiti dei
cluster.
Un cluster intero, da codice
Crea un cluster da una ricerca e scrivi in un file CSV i valori estratti da ogni sito. Le librerie client offrono la stessa cosa sotto forma di funzioni pronte e di riga di comando.
Python
import csv, os, time, requests
KEY = os.environ["PUBLICWWW_KEY"]
BASE = "https://api.publicwww.com"
H = {"Authorization": "Bearer " + KEY}
def call(method, path, body=None):
while True:
r = requests.request(method, BASE + path, headers=H, json=body)
if r.status_code == 429 and r.json()["error"]["code"] == "too_many_requests":
time.sleep(int(r.headers.get("Retry-After", 30)))
continue
r.raise_for_status()
return r.json()
cluster = call("POST", "/v1/clusters", {"query": '"googletagmanager.com/gtm.js"'})
offset = 0
with open("extract.csv", "w", newline="") as f:
out = csv.writer(f)
while offset is not None:
part = call("POST", "/v1/clusters/%d/extract" % cluster["id"],
{"presets": ["gtm", "email"], "offset": offset})
for row in part["rows"]:
out.writerow([row["domain"]] + [" ".join(v) for v in row["values"]])
offset = part["next_offset"]
JavaScript (Node 18+)
const BASE = "https://api.publicwww.com";
const H = { Authorization: "Bearer " + process.env.PUBLICWWW_KEY,
"Content-Type": "application/json" };
async function call(method, path, body) {
for (;;) {
const r = await fetch(BASE + path, { method, headers: H,
body: body && JSON.stringify(body) });
const data = await r.json();
if (r.status === 429 && data.error.code === "too_many_requests") {
await new Promise(ok => setTimeout(ok, 1000 * (r.headers.get("Retry-After") || 30)));
continue;
}
if (!r.ok) throw new Error(data.error.message);
return data;
}
}
const cluster = await call("POST", "/v1/clusters", { query: '"hotjar.com"' });
for (let offset = 0; offset !== null; ) {
const part = await call("POST", `/v1/clusters/${cluster.id}/extract`,
{ presets: ["hotjar", "email"], offset });
for (const row of part.rows) console.log(row.domain, row.values.map(v => v.join(" ")).join(";"));
offset = part.next_offset;
}
PHP
<?php
function call ($method, $path, $body = null) {
$ch = curl_init ("https://api.publicwww.com" . $path);
curl_setopt_array ($ch, [
CURLOPT_CUSTOMREQUEST => $method,
CURLOPT_RETURNTRANSFER => true,
CURLOPT_HTTPHEADER => ["Authorization: Bearer " . getenv ("PUBLICWWW_KEY"),
"Content-Type: application/json"],
CURLOPT_POSTFIELDS => $body === null ? null : json_encode ($body),
]);
$data = json_decode (curl_exec ($ch), true);
if (isset ($data ["error"])) throw new Exception ($data ["error"]["message"]);
return $data;
}
$cluster = call ("POST", "/v1/clusters", ["query" => '"jquery.min.js"']);
$out = fopen ("extract.csv", "w");
for ($offset = 0; $offset !== null; ) {
$part = call ("POST", "/v1/clusters/" . $cluster ["id"] . "/extract",
["presets" => ["email", "phone"], "offset" => $offset]);
foreach ($part ["rows"] as $row)
fputcsv ($out, array_merge ([$row ["domain"]], array_map (fn ($v) => join (" ", $v), $row ["values"])));
$offset = $part ["next_offset"];
}
Go
package main
import (
"bytes"
"encoding/json"
"fmt"
"net/http"
"os"
"strings"
)
func call(method, path string, body, out any) error {
b, _ := json.Marshal(body)
req, _ := http.NewRequest(method, "https://api.publicwww.com"+path, bytes.NewReader(b))
req.Header.Set("Authorization", "Bearer "+os.Getenv("PUBLICWWW_KEY"))
req.Header.Set("Content-Type", "application/json")
resp, err := http.DefaultClient.Do(req)
if err != nil {
return err
}
defer resp.Body.Close()
if resp.StatusCode >= 300 {
return fmt.Errorf("publicwww: %s", resp.Status)
}
return json.NewDecoder(resp.Body).Decode(out)
}
func main() {
var cluster struct{ ID int `json:"id"` }
if err := call("POST", "/v1/clusters", map[string]any{"query": `"googletagmanager.com/gtm.js"`}, &cluster); err != nil {
panic(err)
}
for offset := 0; ; {
var part struct {
Rows []struct {
Domain string `json:"domain"`
Values [][]string `json:"values"`
} `json:"rows"`
NextOffset *int `json:"next_offset"`
}
path := fmt.Sprintf("/v1/clusters/%d/extract", cluster.ID)
if err := call("POST", path, map[string]any{"presets": []string{"gtm", "ga4"}, "offset": offset}, &part); err != nil {
panic(err)
}
for _, r := range part.Rows {
cells := []string{r.Domain}
for _, v := range r.Values {
cells = append(cells, strings.Join(v, " "))
}
fmt.Println(strings.Join(cells, ";"))
}
if part.NextOffset == nil {
break
}
offset = *part.NextOffset
}
}
Ruby
require "json"
require "net/http"
def call(path, body)
uri = URI("https://api.publicwww.com" + path)
req = Net::HTTP::Post.new(uri, "Authorization" => "Bearer #{ENV.fetch('PUBLICWWW_KEY')}",
"Content-Type" => "application/json")
req.body = body.to_json
res = Net::HTTP.start(uri.host, uri.port, use_ssl: true) { |h| h.request(req) }
data = JSON.parse(res.body)
raise data["error"]["message"] if data["error"]
data
end
cluster = call("/v1/clusters", { query: '"hotjar.com"' })
offset = 0
while offset
part = call("/v1/clusters/#{cluster['id']}/extract", { presets: %w[hotjar email], offset: offset })
part["rows"].each { |r| puts [r["domain"], *r["values"].map { |v| v.join(" ") }].join(";") }
offset = part["next_offset"]
end
Errori
| Stato e codice | Significato |
|---|---|
404 cluster_not_found | Nessun cluster con quell'ID in questo account. |
409 cluster_limit | L'account conserva già 100 cluster. |
400 invalid_regex | Un'espressione non è una PCRE tra barre o barre verticali, lunga al massimo 200 caratteri. |
400 unknown_preset | Modello predefinito inesistente; la risposta elenca quelli disponibili. |
400 missing_source, ambiguous_source | Per creare un cluster serve uno solo tra query e domains. |
429 extract_quota_exceeded | I punti di estrazione di oggi sono esauriti. |
L'elenco completo è nella pagina degli errori e nella descrizione che l'API dà di se stessa all'indirizzo https://api.publicwww.com/. In un assistente AI le stesse operazioni sono disponibili come strumenti MCP.
Successivo Esempi di codice