Tutorials

Build a Rich Contact Research List: Laravel + Website to Company Data API

Build a Rich Contact Research List: Laravel + Website to Company Data API

A spreadsheet full of company websites looks like a useful prospecting asset, but it is still several steps away from a usable research list. Someone must visit each site, identify the company, find available contact details, record relevant people, and preserve enough context for a human to review the result.

This tutorial builds that workflow in Laravel. A console command imports a CSV file, queue jobs enrich each website through the Website to Company data API, and a second command exports a reviewable CSV containing company, contact, email, phone, and people data. The design stays deliberately modest: Laravel’s built-in HTTP client, database queue, structured application-boundary mapping, and no unnecessary packages.

Get access and copy the service token

Before writing integration code, create or access your account. Register here, or use the sign-in page if you already have an account.

  1. Open the Website to Company data service page.
  2. Choose the available Free, Plus, or Pro plan and complete its activation.
  3. Open the official service documentation.
  4. Find the Service token panel and copy the service-scoped token.

This service requires authentication. Its token belongs in the token query parameter. Regenerating the token revokes the previously active token, so coordinate rotation with deployment and update every environment that uses it.

The exact request is GET https://ai.mihajlo.mk/api/website-to-company-data/v1/extract. Verify access with a public company website before building the application:

curl --get \
  'https://ai.mihajlo.mk/api/website-to-company-data/v1/extract' \
  --data-urlencode 'website=https://example.com' \
  --data-urlencode 'token=YOUR_SERVICE_TOKEN'

Do not paste a real token into source control, shell history shared with others, logs, screenshots, or fixtures. Put it in the Laravel environment file, and add only the placeholder to .env.example:

# .env
WEBSITE_TO_COMPANY_TOKEN=YOUR_SERVICE_TOKEN
QUEUE_CONNECTION=database

# .env.example
WEBSITE_TO_COMPANY_TOKEN=

Choose a small, resilient architecture

The input spreadsheet should be exported as CSV with a website header. CSV avoids coupling the application to a particular office format. Each normalized website becomes a database record, and one queued job performs its enrichment. The database record is both durable workflow state and the source for the final export.

Background jobs are worthwhile here because network calls are slower and less predictable than CSV parsing. They also let operators control throughput by changing worker count without redesigning the importer. The trade-off is eventual completion: users import first, monitor the queue, and export only after jobs settle.

The relevant project structure is:

app/
  Console/Commands/ImportCompanyResearch.php
  Console/Commands/ExportCompanyResearch.php
  Data/CompanyResearchData.php
  Exceptions/WebsiteDataFailure.php
  Jobs/EnrichCompanyWebsite.php
  Models/CompanyResearchItem.php
  Services/WebsiteToCompanyClient.php
config/services.php
database/migrations/..._create_company_research_items_table.php
tests/Feature/EnrichCompanyWebsiteTest.php
tests/Unit/WebsiteToCompanyClientTest.php

Create a Laravel application configured for PHP 8.3 or later, then generate the classes with normal Artisan commands. If the application does not already contain a jobs-table migration for the database queue, generate one before migrating.

php artisan make:model CompanyResearchItem -m
php artisan make:job EnrichCompanyWebsite
php artisan make:command ImportCompanyResearch
php artisan make:command ExportCompanyResearch
php artisan make:test WebsiteToCompanyClientTest --unit
php artisan make:test EnrichCompanyWebsiteTest
php artisan make:queue-table
php artisan migrate

Persist workflow state and API fields

The migration stores every required response area as JSON because the service may return structured values rather than a single string. It also keeps failures separate from successful data, making partial runs easy to inspect.

<?php

use Illuminate\Database\Migrations\Migration;
use Illuminate\Database\Schema\Blueprint;
use Illuminate\Support\Facades\Schema;

return new class extends Migration {
    public function up(): void
    {
        Schema::create('company_research_items', function (Blueprint $table) {
            $table->id();
            $table->string('source_url')->unique();
            $table->string('status')->default('pending')->index();
            $table->json('company')->nullable();
            $table->json('contact')->nullable();
            $table->json('email')->nullable();
            $table->json('phone')->nullable();
            $table->json('people')->nullable();
            $table->string('failure_code')->nullable();
            $table->text('failure_message')->nullable();
            $table->timestamp('reviewed_at')->nullable();
            $table->timestamps();
        });
    }

    public function down(): void
    {
        Schema::dropIfExists('company_research_items');
    }
};

In CompanyResearchItem, make these columns fillable and cast the five API fields to arrays. Cast reviewed_at to datetime. Keeping raw-but-validated structures preserves useful service detail without leaking transport concerns throughout the application.

<?php

namespace App\Models;

use Illuminate\Database\Eloquent\Model;

final class CompanyResearchItem extends Model
{
    protected $fillable = [
        'source_url', 'status', 'company', 'contact', 'email',
        'phone', 'people', 'failure_code', 'failure_message',
        'reviewed_at',
    ];

    protected function casts(): array
    {
        return [
            'company' => 'array',
            'contact' => 'array',
            'email' => 'array',
            'phone' => 'array',
            'people' => 'array',
            'reviewed_at' => 'datetime',
        ];
    }
}

Map uncertain data at the boundary

Remote JSON should not be trusted merely because it decoded successfully. The mapper below accepts only the documented top-level areas: company, contact, email, phone, and people. Arrays are retained; scalar values are wrapped consistently; objects or resources cannot enter the domain model.

<?php

namespace App\Data;

use UnexpectedValueException;

final readonly class CompanyResearchData
{
    public function __construct(
        public ?array $company,
        public ?array $contact,
        public ?array $email,
        public ?array $phone,
        public ?array $people,
    ) {}

    public static function fromResponse(array $payload): self
    {
        $keys = ['company', 'contact', 'email', 'phone', 'people'];

        if (! array_any($keys, fn (string $key) => array_key_exists($key, $payload))) {
            throw new UnexpectedValueException('The response contains no recognized data fields.');
        }

        $field = static function (string $key) use ($payload): ?array {
            if (! array_key_exists($key, $payload) || $payload[$key] === null) {
                return null;
            }

            if (is_array($payload[$key])) {
                return $payload[$key];
            }

            if (is_scalar($payload[$key])) {
                return ['value' => (string) $payload[$key]];
            }

            throw new UnexpectedValueException("Invalid {$key} field.");
        };

        return new self(...array_map($field, $keys));
    }
}

array_any is available in PHP 8.4, not PHP 8.3, so a PHP 8.3 project should replace that condition with a small loop. This keeps the tutorial’s stated runtime honest:

$recognized = false;

foreach ($keys as $key) {
    if (array_key_exists($key, $payload)) {
        $recognized = true;
        break;
    }
}

if (! $recognized) {
    throw new UnexpectedValueException(
        'The response contains no recognized data fields.'
    );
}

Build the bounded HTTP client

Add the service configuration to the array returned by config/services.php:

'website_to_company' => [
    'token' => env('WEBSITE_TO_COMPANY_TOKEN'),
],

The client uses short connection and response timeouts. It retries connection failures, HTTP 429, and server failures, but never retries authentication or validation responses. Backoff is bounded, and a numeric Retry-After value is honored up to five seconds.

<?php

namespace App\Services;

use App\Data\CompanyResearchData;
use App\Exceptions\WebsiteDataFailure;
use Illuminate\Http\Client\ConnectionException;
use Illuminate\Support\Facades\Http;
use Throwable;

final class WebsiteToCompanyClient
{
    private const ENDPOINT =
        'https://ai.mihajlo.mk/api/website-to-company-data/v1/extract';

    public function extract(string $website): CompanyResearchData
    {
        $token = config('services.website_to_company.token');

        if (! is_string($token) || $token === '') {
            throw new WebsiteDataFailure('configuration', 'Service token is missing.');
        }

        for ($attempt = 1; $attempt <= 3; $attempt++) {
            try {
                $response = Http::acceptJson()
                    ->connectTimeout(3)
                    ->timeout(15)
                    ->get(self::ENDPOINT, [
                        'website' => $website,
                        'token' => $token,
                    ]);
            } catch (ConnectionException $exception) {
                if ($attempt === 3) {
                    throw new WebsiteDataFailure(
                        'connection',
                        'The service could not be reached.',
                        $exception
                    );
                }

                $this->pause($attempt);
                continue;
            }

            if ($response->successful()) {
                try {
                    return CompanyResearchData::fromResponse($response->json());
                } catch (Throwable $exception) {
                    throw new WebsiteDataFailure(
                        'invalid_response',
                        'The service returned an unusable response.',
                        $exception
                    );
                }
            }

            if (in_array($response->status(), [401, 403], true)) {
                throw new WebsiteDataFailure('authentication', 'Service authentication failed.');
            }

            if (in_array($response->status(), [400, 404, 422], true)) {
                throw new WebsiteDataFailure('validation', 'The website was rejected.');
            }

            if ($response->status() === 429 || $response->serverError()) {
                if ($attempt < 3) {
                    $this->pause($attempt, $response->header('Retry-After'));
                    continue;
                }

                $code = $response->status() === 429 ? 'rate_limit' : 'upstream';
                throw new WebsiteDataFailure($code, 'The service is temporarily unavailable.');
            }

            throw new WebsiteDataFailure(
                'http_error',
                'Unexpected service response: '.$response->status()
            );
        }

        throw new WebsiteDataFailure('unknown', 'Enrichment did not complete.');
    }

    private function pause(int $attempt, ?string $retryAfter = null): void
    {
        $milliseconds = ctype_digit((string) $retryAfter)
            ? min(5000, (int) $retryAfter * 1000)
            : 250 * (2 ** ($attempt - 1));

        usleep($milliseconds * 1000);
    }
}

WebsiteDataFailure is a small custom exception with a public string kind property and an optional previous exception. Its messages are deliberately safe: neither the query URL nor token is copied into logs or database rows.

Import, enrich, and export

The import command should read the header with fgetcsv, locate the website column, and reject malformed rows. Accept only http and https, require a hostname, and reject credentials, IP literals, localhost, fragments, and nonstandard ports. Normalize ordinary company inputs to lowercase scheme and hostname before using firstOrCreate.

Dispatch one EnrichCompanyWebsite job for each new or previously failed record. Add a --refresh option for intentional re-enrichment; otherwise, completed and active records should be skipped to prevent duplicate API use.

<?php

namespace App\Jobs;

use App\Exceptions\WebsiteDataFailure;
use App\Models\CompanyResearchItem;
use App\Services\WebsiteToCompanyClient;
use Illuminate\Contracts\Queue\ShouldQueue;
use Illuminate\Foundation\Queue\Queueable;
use Illuminate\Support\Facades\Log;
use Throwable;

final class EnrichCompanyWebsite implements ShouldQueue
{
    use Queueable;

    public int $tries = 1;
    public int $timeout = 40;

    public function __construct(public int $itemId)
    {
        $this->onQueue('research');
    }

    public function handle(WebsiteToCompanyClient $client): void
    {
        $item = CompanyResearchItem::findOrFail($this->itemId);
        $item->update(['status' => 'processing']);

        $started = hrtime(true);

        try {
            $data = $client->extract($item->source_url);

            $item->update([
                'status' => 'completed',
                'company' => $data->company,
                'contact' => $data->contact,
                'email' => $data->email,
                'phone' => $data->phone,
                'people' => $data->people,
                'failure_code' => null,
                'failure_message' => null,
            ]);

            Log::info('Company research completed', [
                'item_id' => $item->id,
                'duration_ms' => (int) ((hrtime(true) - $started) / 1_000_000),
            ]);
        } catch (WebsiteDataFailure $exception) {
            $item->update([
                'status' => 'failed',
                'failure_code' => $exception->kind,
                'failure_message' => $exception->getMessage(),
            ]);

            Log::warning('Company research failed', [
                'item_id' => $item->id,
                'failure_code' => $exception->kind,
            ]);
        }
    }

    public function failed(?Throwable $exception): void
    {
        CompanyResearchItem::whereKey($this->itemId)->update([
            'status' => 'failed',
            'failure_code' => 'job_failure',
            'failure_message' => 'The background job failed unexpectedly.',
        ]);
    }
}

The export command should write to Laravel’s local storage disk using fputcsv. Emit the columns source_url, status, company, contact, email, phone, people, and failure_code. Encode each structured field with json_encode. That produces a spreadsheet-friendly audit artifact without flattening or silently discarding nested data.

php artisan research:import storage/app/imports/websites.csv
php artisan queue:work --queue=research --sleep=1 --tries=1 --timeout=45
php artisan research:export company-research.csv

The resulting file lives under storage/app. Reviewers can filter failures, inspect structured cells, correct source websites, and mark accepted rows through a later application workflow using reviewed_at.

Test the boundary and the job

Use Http::fake() so tests never consume quota or depend on network availability. One unit test should return representative company, contact, email, phone, and people values, assert their mapping, and inspect the outgoing query to confirm both required parameters. A second test should return 401, assert an authentication failure, and verify that only one request occurred.

public function test_it_maps_the_service_response(): void
{
    config(['services.website_to_company.token' => 'test-token']);

    Http::fake([
        'https://ai.mihajlo.mk/api/website-to-company-data/v1/extract*' =>
            Http::response([
                'company' => ['name' => 'Example Company'],
                'contact' => ['page' => '/contact'],
                'email' => ['[email protected]'],
                'phone' => ['+1 555 0100'],
                'people' => [['name' => 'Alex Example']],
            ], 200),
    ]);

    $data = app(WebsiteToCompanyClient::class)->extract('https://example.com');

    $this->assertSame('Example Company', $data->company['name']);

    Http::assertSent(function (Request $request): bool {
        parse_str(parse_url($request->url(), PHP_URL_QUERY), $query);

        return $request->method() === 'GET'
            && $query['website'] === 'https://example.com'
            && $query['token'] === 'test-token';
    });
}

The feature test should create a pending item, fake a successful response, invoke the job’s handle method through the container, and assert that the database row is completed with mapped JSON. Also cover a 422 response becoming a failed row with validation. These tests exercise transport, mapping, and persistence without testing Laravel itself.

Operate it safely in production

Run the queue worker under the process manager already used by the deployment platform, and restart workers during deployment so they load new code. Keep worker concurrency conservative until the plan’s quota and rate behavior are understood. The client’s three internal attempts are the only automatic retries; setting the job to one attempt prevents nested retry storms.

Cache configuration only after the production token is present with php artisan config:cache. Rotate the token by updating the environment, rebuilding the configuration cache, and restarting workers. Because regenerating the service token immediately revokes its predecessor, perform those steps as one controlled change.

Monitor counts of pending, processing, completed, and failed records. Alert on sustained authentication, rate_limit, or invalid_response failures. Avoid logging response bodies: contact research can contain personal data, and operational logs usually have broader access and longer retention than the application database.

Common failures

  • Every request fails authentication: verify activation, the service-scoped token, configuration caching, and whether someone regenerated the token.
  • Rows remain pending: confirm a worker is running on the research queue and that its database connection can see the jobs table.
  • Rows remain processing: inspect failed jobs and worker termination events; safely requeue affected records through the explicit refresh path.
  • Many rate-limit failures appear: reduce worker concurrency and retry failed rows later instead of extending unbounded sleeps.
  • The response is marked invalid: retain the failure classification, compare the current official documentation, and update only the boundary mapper.

Final verification checklist

  • The account and Free, Plus, or Pro plan are active.
  • The service token exists only in environment-backed configuration.
  • The minimal GET request succeeds with website and token.
  • The CSV header contains website, and unsafe URLs are rejected.
  • Queue workers consume the research queue with bounded concurrency.
  • Company, contact, email, phone, and people data are mapped before persistence.
  • Authentication and validation failures are not retried.
  • Rate limits, connection failures, and server errors receive bounded backoff.
  • Tests use Http::fake() and contain no real credential.
  • The exported CSV contains completed rows and visible failure states for review.

The important result is not merely a working API request. It is a controlled research pipeline: inputs are validated, remote behavior is contained, failures remain legible, credentials stay protected, and every enriched row returns to a human-readable spreadsheet. That is the difference between a clever script and a production integration a small team can safely keep using.

Blog author portrait

Mihajlo

I’m Mihajlo — a developer driven by curiosity, discipline, and the constant urge to create something meaningful. I share insights, tutorials, and free services to help others simplify their work and grow in the ever-evolving world of software and AI.