Casos de Uso

Tratamento de CAPTCHA para coleta de dados de salários e remunerações

Bancos de dados de salários, painéis de empregos e portais trabalhistas governamentais protegem os dados de remuneração com Cloudflare Turnstile e reCAPTCHA. Os CAPTCHAs são acionados ao consultar faixas salariais por função, localização ou setor, especialmente durante a coleta de dados em massa de vários cargos. Veja como lidar com eles.

Padrões CAPTCHA em portais salariais

Tipo de fonte CAPTCHA Gatilho
Sites de comparação de salários Cloudflare Turnstile Consultas de pesquisa repetidas
Filtros salariais do quadro de empregos reCAPTCHA v2 Várias pesquisas salariais
Estatísticas trabalhistas do governo Imagem CAPTCHA Solicitações de download de dados
Páginas salariais corporativas Cloudflare Turnstile em staging Visualizações de página em massa
Plataformas de pesquisa de RH reCAPTCHA v3 Envios de formulário

Coletor de dados salariais

import requests
import time
import re
from dataclasses import dataclass

@dataclass
class SalaryRecord:
    title: str
    location: str
    min_salary: float
    max_salary: float
    median_salary: float
    sample_size: int
    source: str

class SalaryCollector:
    def __init__(self, api_key):
        self.api_key = api_key
        self.session = requests.Session()
        self.session.headers.update({
            "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36"
        })

    def collect_salary_data(self, portal_url, job_title, location):
        """Search for salary data, solving CAPTCHAs as needed."""
        response = self.session.get(portal_url, params={
            "title": job_title,
            "location": location
        })

        if self._is_turnstile_challenge(response):
            response = self._solve_turnstile_and_retry(response, portal_url)

        return self._parse_salary_data(response.text, portal_url)

    def collect_bulk(self, portal_url, job_titles, locations):
        """Collect salary data for multiple job title + location combos."""
        results = []

        for title in job_titles:
            for location in locations:
                try:
                    data = self.collect_salary_data(
                        portal_url, title, location
                    )
                    results.extend(data)
                    # Respectful delay between requests
                    time.sleep(2)
                except Exception as e:
                    print(f"Failed for {title} in {location}: {e}")

        return results

    def _is_turnstile_challenge(self, response):
        return (
            response.status_code == 403 or
            "cf-turnstile" in response.text or
            "challenges.cloudflare.com" in response.text
        )

    def _solve_turnstile_and_retry(self, response, url):
        match = re.search(r'data-sitekey="(0x[^"]+)"', response.text)
        if not match:
            raise ValueError("Turnstile sitekey not found")

        resp = requests.post("https://ocr.captchaai.com/in.php", data={
            "key": self.api_key,
            "method": "turnstile",
            "sitekey": match.group(1),
            "pageurl": url,
            "json": 1
        })
        task_id = resp.json()["request"]

        for _ in range(60):
            time.sleep(3)
            result = requests.get("https://ocr.captchaai.com/res.php", params={
                "key": self.api_key,
                "action": "get",
                "id": task_id,
                "json": 1
            })
            data = result.json()
            if data["status"] == 1:
                return self.session.post(url, data={
                    "cf-turnstile-response": data["request"]
                })

        raise TimeoutError("Turnstile solve timed out")

    def _parse_salary_data(self, html, source):
        from bs4 import BeautifulSoup
        soup = BeautifulSoup(html, "html.parser")
        records = []

        def text_or_empty(node):
            return node.text.strip() if node and node.text else ""

        for row in soup.select(".salary-row, .compensation-entry, tr[data-salary]"):
            try:
                records.append(SalaryRecord(
                    title=text_or_empty(row.select_one(".job-title, .title")),
                    location=text_or_empty(row.select_one(".location")),
                    min_salary=self._parse_amount(
                        text_or_empty(row.select_one(".min-salary, .low"))
                    ),
                    max_salary=self._parse_amount(
                        text_or_empty(row.select_one(".max-salary, .high"))
                    ),
                    median_salary=self._parse_amount(
                        text_or_empty(row.select_one(".median, .mid"))
                    ),
                    sample_size=int(
                        text_or_empty(row.select_one(".count, .sample")).replace(",", "") or 0
                    ),
                    source=source
                ))
            except (AttributeError, ValueError):
                continue

        return records

    def _parse_amount(self, text):
        if not text:
            return 0.0
        cleaned = re.sub(r'[^\d.]', '', text)
        return float(cleaned) if cleaned else 0.0


# Usage
collector = SalaryCollector("YOUR_API_KEY")
data = collector.collect_bulk(
    "https://salary.example.com/search",
    job_titles=["Software Engineer", "Data Analyst", "Product Manager"],
    locations=["San Francisco", "New York", "Austin"]
)

for record in data:
    print(f"{record.title} in {record.location}: "
          f"${record.min_salary:,.0f}–${record.max_salary:,.0f} "
          f"(median: ${record.median_salary:,.0f})")

Agregação de múltiplas fontes (JavaScript)

class SalaryAggregator {
  constructor(apiKey) {
    this.apiKey = apiKey;
    this.sources = [];
  }

  addSource(name, searchUrl) {
    this.sources.push({ name, searchUrl });
  }

  async collectForRole(jobTitle, location) {
    const results = [];

    for (const source of this.sources) {
      try {
        const data = await this.querySource(source, jobTitle, location);
        results.push({ source: source.name, ...data });
      } catch (error) {
        results.push({ source: source.name, error: error.message });
      }
    }

    return this.aggregateResults(results, jobTitle, location);
  }

  async querySource(source, jobTitle, location) {
    const url = `${source.searchUrl}?title=${encodeURIComponent(jobTitle)}&location=${encodeURIComponent(location)}`;
    const response = await fetch(url);
    const html = await response.text();

    if (html.includes('cf-turnstile') || response.status === 403) {
      return this.solveAndRetry(source.searchUrl, html, jobTitle, location);
    }

    return this.parseSalaryData(html);
  }

  async solveAndRetry(baseUrl, html, jobTitle, location) {
    const match = html.match(/data-sitekey="(0x[^"]+)"/);
    if (!match) throw new Error('Turnstile sitekey not found');

    const submitResp = await fetch('https://ocr.captchaai.com/in.php', {
      method: 'POST',
      body: new URLSearchParams({
        key: this.apiKey,
        method: 'turnstile',
        sitekey: match[1],
        pageurl: baseUrl,
        json: '1'
      })
    });
    const { request: taskId } = await submitResp.json();

    for (let i = 0; i < 60; i++) {
      await new Promise(r => setTimeout(r, 3000));
      const result = await fetch(
        `https://ocr.captchaai.com/res.php?key=${this.apiKey}&action=get&id=${taskId}&json=1`
      );
      const data = await result.json();
      if (data.status === 1) {
        const response = await fetch(baseUrl, {
          method: 'POST',
          body: new URLSearchParams({
            'cf-turnstile-response': data.request,
            title: jobTitle,
            location: location
          })
        });
        return this.parseSalaryData(await response.text());
      }
    }
    throw new Error('Turnstile solve timed out');
  }

  aggregateResults(results, jobTitle, location) {
    const valid = results.filter(r => !r.error && r.median);
    if (valid.length === 0) return null;

    const medians = valid.map(r => r.median);
    return {
      jobTitle,
      location,
      avgMedian: medians.reduce((a, b) => a + b, 0) / medians.length,
      sources: valid.length,
      range: { min: Math.min(...medians), max: Math.max(...medians) }
    };
  }
}

// Usage
const aggregator = new SalaryAggregator('YOUR_API_KEY');
aggregator.addSource('SalaryDB', 'https://salarydb.example.com/search');
aggregator.addSource('PayScale', 'https://payscale.example.com/lookup');

const result = await aggregator.collectForRole('Software Engineer', 'San Francisco');
console.log(`Median salary: $${result.avgMedian.toLocaleString()} (${result.sources} sources)`);

Estratégia de coleta de dados

Abordagem Volume por dia Frequência CAPTCHA mais adequado para
Sequencial com atrasos 100 a 500 consultas Baixo Pequenas pesquisas
Rotação de proxy 500–2.000 consultas Moderado Análise regional
Paralelo multi-sessão 2.000 a 10.000 consultas Alto Conjuntos de dados abrangentes

Solução de problemas

Problema Causa Correção
Cloudflare Turnstile em todas as pesquisas Sessão expirada Persistir cookie cookie_qa_validacao
Os dados salariais mostram "Login obrigatório" Portal requer autenticação Autentique antes de pesquisar
Resultados vazios após resolução do CAPTCHA Parâmetros POST ausentes Incluir todos os campos ocultos do formulário
Dados inconsistentes entre execuções Portal mostra diferentes gamas Use parâmetros de consulta consistentes

Perguntas frequentes

Quantas consultas salariais posso fazer por dia?

Depende dos limites de taxa do portal, não de CaptchaAI. CaptchaAI resolve Cloudflare Turnstile com 100% de sucesso. Espace as solicitações com intervalo de 2 a 5 segundos e gire os proxies para coleta de alto volume.

Devo usar proxies para coleta de dados salariais?

Sim, especialmente para coleta em massa de milhares de cargos. Os proxies residenciais reduzem significativamente a frequência CAPTCHA em comparação com IPs de datacenters.

Posso coletar dados salariais em tempo real?

A maioria dos portais salariais atualiza os dados mensalmente ou trimestralmente, portanto, a coleta em tempo real é desnecessária. Agende execuções de coleta semanais ou mensais para conjuntos de dados abrangentes.

Artigos relacionados

Próximas etapas

Colete dados de remuneração de forma confiável -obtenha sua chave API CaptchaAIe gerenciar CAPTCHAs do portal de salários automaticamente.

Os comentários estão desativados para este artigo.