GPUnex
GPU Rental 6 min read

API Reference

Manage your GPU instances programmatically. Learn how to authenticate, create instances, check status, and more via the GPUnex API.

Getting Your API Key

The GPUnex API allows you to manage GPU instances, check balances, and automate your workflows without using the web dashboard. Before making any API requests, you need to generate an API key.

  1. Log in to your GPUnex account and navigate to your Profile page or the Dashboard settings panel.

  2. Open the API Keys section. Look for the API Keys tab or card within your account settings.

  3. Click “Create New API Key”. You will be prompted to give the key a descriptive name. Choose something meaningful that identifies the key’s purpose — for example, “Production Server”, “CI/CD Pipeline”, or “Local Development”. This makes it easier to manage multiple keys later.

  4. Copy your API key immediately. Once the key is generated, it will be displayed exactly once. Copy it and store it in a secure location such as a password manager or an encrypted secrets vault. You will not be able to view the full key again after this step.

  5. Understand key scoping and revocation. Each API key is scoped to your account and inherits your account permissions. You can create multiple keys for different applications or environments. If a key is compromised or no longer needed, you can revoke it at any time from the API Keys section. Revoking a key is immediate and permanent — all requests using that key will be rejected.

Authentication

All requests to the GPUnex API must include an Authorization header with your API key using the Bearer token scheme.

Header format:

Authorization: Bearer YOUR_API_KEY
Terminal
$

If the API key is missing, invalid, or revoked, the API will return a 401 Unauthorized response:

{
  "error": "unauthorized",
  "message": "Invalid or missing API key. Please check your Authorization header."
}

Make sure your key is included in every request. The API does not support session-based authentication or cookie-based authentication.

Core Endpoints

The GPUnex API is organized around RESTful resources. All endpoints use the base URL https://api.gpunex.com/v1. Below is a summary of the core endpoints available.

MethodEndpointDescription
GET/v1/instancesList all your active and recent instances. Returns an array of instance objects with their current status, GPU model, and configuration details.
POST/v1/instancesCreate a new GPU instance. Requires a JSON body specifying the GPU model, framework, and region. Charges are deducted from your USDC balance.
GET/v1/instances/:idGet detailed information about a specific instance by its unique ID. Includes SSH connection details, runtime metrics, and billing information.
DELETE/v1/instances/:idTerminate a running instance. The instance will be stopped and you will no longer be charged for it. Any unsaved data on the instance will be lost.
GET/v1/gpu-modelsList all GPU models currently available on the marketplace. Includes pricing, VRAM, availability by region, and supported frameworks.
GET/v1/balanceCheck your current USDC wallet balance and recent transaction summary.

All responses are returned in JSON format. Successful requests return a 200 OK status code for GET requests and a 201 Created status code for POST requests that create resources.

Example: Creating an Instance

To create a new GPU instance, send a POST request to /v1/instances with a JSON body specifying your desired configuration.

Request:

Terminal
$

Request body parameters:

ParameterTypeRequiredDescription
gpu_modelstringYesThe GPU model identifier. Use /v1/gpu-models to see available options (e.g., H100_80GB, A100_80GB, L40S_48GB, L4_24GB).
frameworkstringYesThe pre-installed framework and version (e.g., pytorch-2.3, tensorflow-2.16, jax-0.4).
regionstringYesThe datacenter region for the instance (e.g., us-east-1, eu-west-1, ap-southeast-1).

Example response (201 Created):

{
  "id": "inst_7f3a9b2c4d1e",
  "status": "provisioning",
  "gpu_model": "A100_80GB",
  "framework": "pytorch-2.3",
  "region": "us-east-1",
  "ssh_command": null,
  "hourly_rate": "1.89",
  "currency": "USDC",
  "created_at": "2026-02-15T14:32:07Z"
}

The instance will initially have a status of provisioning. Once the GPU is allocated and the environment is ready, the status will change to running and the ssh_command field will be populated with your connection string. Provisioning typically takes between 30 seconds and 2 minutes depending on availability.

Example: Checking Instance Status

Once you have created an instance, you can check its current status at any time by querying the instance endpoint with its ID.

Request:

curl -H "Authorization: Bearer YOUR_API_KEY" \
  https://api.gpunex.com/v1/instances/inst_7f3a9b2c4d1e

Example response (200 OK):

{
  "id": "inst_7f3a9b2c4d1e",
  "status": "running",
  "gpu_model": "A100_80GB",
  "framework": "pytorch-2.3",
  "region": "us-east-1",
  "ssh_command": "ssh [email protected] -p 2222",
  "ip_address": "203.0.113.42",
  "hourly_rate": "1.89",
  "currency": "USDC",
  "uptime_seconds": 3847,
  "total_cost": "2.02",
  "created_at": "2026-02-15T14:32:07Z",
  "started_at": "2026-02-15T14:33:15Z"
}

Status values:

StatusMeaning
provisioningThe instance is being set up. GPU resources are being allocated and the framework environment is being prepared.
runningThe instance is active and ready for use. SSH access is available.
stoppingThe instance is in the process of shutting down.
terminatedThe instance has been stopped and is no longer incurring charges.
errorAn error occurred during provisioning or runtime. Contact support if this persists.

Rate Limits and Best Practices

Rate Limits

The GPUnex API enforces rate limits to ensure fair usage and platform stability. The current limits are:

  • General endpoints: 120 requests per minute per API key.
  • Instance creation: 10 requests per minute per API key.
  • Balance and read-only endpoints: 300 requests per minute per API key.

If you exceed the rate limit, the API will return a 429 Too Many Requests response with a Retry-After header indicating how many seconds to wait before retrying.

{
  "error": "rate_limit_exceeded",
  "message": "Too many requests. Please retry after 12 seconds.",
  "retry_after": 12
}

Best Practices

Follow these guidelines to keep your integration secure and reliable.

  1. Never expose API keys in client-side code. Do not embed your API key in JavaScript running in the browser, mobile app source code, or any publicly accessible repository. API keys should only be used in server-side applications where they cannot be inspected by end users.

Important

Never expose your API key in client-side code, browser JavaScript, or public repositories. API keys should only be used in server-side applications.

  1. Use environment variables. Store your API key in an environment variable rather than hardcoding it in your source files. For example:

    export GPUNEX_API_KEY="your_api_key_here"

    Then reference it in your code:

    curl -H "Authorization: Bearer $GPUNEX_API_KEY" https://api.gpunex.com/v1/instances
  2. Rotate keys periodically. As a security best practice, generate a new API key every 90 days and revoke the old one. This limits the impact if a key is inadvertently leaked.

  3. Use separate keys for separate environments. Create distinct API keys for development, staging, and production. This way, revoking one key does not affect other environments.

Tip

Use separate API keys for development, staging, and production. This way, revoking one key does not affect other environments.

  1. Handle errors gracefully. Always check HTTP status codes in your application. Implement retry logic with exponential backoff for 429 and 5xx responses. Do not retry 4xx errors other than 429 — these indicate a problem with the request itself.

  2. Monitor your usage. Keep track of your API call volume and instance spending. Use the /v1/balance endpoint to programmatically monitor your USDC balance and set up alerts if it falls below a threshold.