Air Assistant, credit breakdown and expiry, payment improvements, GPU auto-upgrade, and replica timeline
Air Assistant agent (beta)Ask about the official documentation and your own account in plain language, right inside the AirCloud console. Open it with the Air Assistant button in the console header. See the Air Assistant guide for details.- Documentation help: Ask about AirCloud features, how-tos, Autoscaling, Air API, and more. Answers are grounded in the official docs and come with links to the relevant pages.
- Account lookups: Check available GPUs and their prices, endpoint and replica status, usage, credits, and invoices.
- Scoped to your access: Only organizations and projects the signed-in user can access are queried, and secrets such as API keys are never shown.
- Read-only beta: For now it only answers questions and explains how to use the platform. Creating, stopping, editing, or deleting resources still happens in the console.
- Suggested and follow-up questions: Suggested questions help you get started, and related follow-ups appear below each answer.
- Conversation controls: Start a new conversation or download the current one as a file. Reopening past conversations is not supported yet.
- Data handling: Questions and answers are kept for 30 days by default and processed by Air API models served by AirCloud. Avoid entering sensitive information.
- Double-check what matters: AI answers can be wrong. Verify costs, payments, and resource status on the corresponding console screens.
- Credit breakdown: Credit breakdown on the Billing page lists your remaining credits by source (payments, sign-up credit, promotions, rewards, and more) and whether each is paid or free.
- Expiry dates: Each credit shows its expiry date and days remaining. Free credits and credits closest to expiring are used first.
- Expiry reminders: You get a notification when some of your credits are about to expire.
- Promo code deadlines: Promo codes can now have a redemption deadline, and credits granted by a code can carry their own validity period.
- Exact eligibility: Whether a payment can be cancelled is now decided by whether the credits granted by that payment are still untouched, not by the organization’s total balance. A payment whose credits were used even partially cannot be cancelled; an unused one can, regardless of how long ago it was made. The 7-day limit is gone.
- Auto top-up included: Auto top-up charges made with a saved card can be cancelled when unused. Recurring contracts settled offline are not covered.
- Before deleting an organization: The organization deletion dialog (including a last owner’s account withdrawal) shows how many of your own card payments can still be cancelled, with a direct link to Payment History.
- Clear failure reasons: When a card payment fails, the console tells you why — issuer decline, additional authentication required, overseas payments blocked — and what to try next.
- Card receipt link: The payment confirmation email links straight to the card issuer’s receipt.
- Extra invoice recipients: Postpaid organizations can send invoices to addresses outside the organization, such as an accounting contact.
- On (default): Your GPU is tried first; if none is free, the endpoint starts right away on an equal-or-better GPU in the same family. You are still billed at the rate of the GPU you selected.
- Off: The endpoint runs only on exactly the GPU you selected and waits until one frees up. Useful for benchmarks and other work where the GPU model affects results.
- Where to set it: In the container creation flow and under Resources in container settings. If your GPU is sold out, you can turn auto-upgrade on right from the creation screen.
- Active replicas card: At the top of the runtime overview, one card shows how the replica count changed and when each replica was running. Replicas that ran only briefly appear as dots.
- Min and max guides: The autoscaling minimum and maximum replica counts are drawn as dashed lines, so you can see how close you are to the limit. Containers with autoscaling off show their fixed replica count instead. Hover ⓘ to see how replicas are scaled.
- Scale markers: Every metric chart marks the moments the replica count changed with a vertical dashed line, so you can compare how requests, latency, GPU utilization, and more changed before and after each scale.
- Drag to zoom: Drag across any chart to zoom every chart on the page into that range, re-queried at a finer interval. Click Reset zoom to go back to the previous range.
- Default metric order: Requests, errors, and latency now come first by default. Any order you arranged yourself is kept.
- Your time zone: Dates and times in the console follow your device’s time zone, and emails and billing documents label the time zone they use.
- Environment variable placeholders: When you create a container,
{endpoint_id},{endpoint_name}, and{project_id}in environment variable values are replaced with the actual values. The container name is fixed at creation time. - GPU locked by templates: When a template specifies an instance type, the GPU selection is locked to that type when you create a container from it.
- Volume mounts carried over: Creating a container from an existing one now carries over its volume mount settings too.
- USD organizations: Organizations paying in USD see balances, rates, payments, and exports in a single consistent unit.
Long-term rentals, in-console requests, unified search, and account controls
Long-term rentalsGPU reservations are now called long-term rentals, and the whole flow lives in the console.- From one month: The minimum term dropped from three months to one, with presets up to a year and a lower hourly rate the longer you rent.
- Priced before you commit: The rental screen shows the hourly rate, the discount applied, and the total for the term before you submit.
- Inquiry in the console: Send a rental inquiry from the same page, track it as a request, and withdraw it if plans change. Our team replies within one business day.
- Rental notifications: Get notified when a rental completes, starts, is about to expire, and expires. Start and expiry also arrive by email.
- Clearer status: Rentals are listed as Requested, Scheduled, Active, Expired, or Cancelled, with days remaining.
- Request a GPU, template, or model: Each picker now has an entry point for “what I need isn’t here”.
- My requests: Every request — waitlists, limit increases, and catalog requests — is tracked in one place, with our team’s reply attached.
- Outcome notifications: You are notified when a request is received and when it is handled.
- One search box: Find containers, projects, models, keys, members, rentals, and requests from a single search in the header.
- Documentation included: Search results reach into the documentation site, linking straight to the relevant section.
- Run actions: Start common actions — create a container, invite a member, issue an API key — without navigating there first.
- Change your password: Set a new password directly in Account Settings. All devices sign out afterward.
- Notification settings: A settings tab lets you choose exactly which events reach the console, your desktop, and your inbox, with a guide to what each notification means.
- Delete an organization or withdraw: Both are now self-service, with an up-front check of what has to be settled first and what will be irrecoverably removed.
- A place to start: Users who don’t belong to an organization yet land on a start screen instead of an empty console.
- Mount a volume at several paths: One volume can be mounted at more than one path inside the container.
- Templates, regrouped: The template picker is organized by what you are trying to do — the workload, how you reach it, and what varies inside.
- Live status: Container status updates on its own, without refreshing or polling.
GPU availability, create from existing container, revamped usage & billing, and expanded notifications
GPU availability and resource requestsCheck GPU stock before you deploy, and request more resources right from the console when you hit a limit.- Availability at a glance: The GPU list in the container create flow now shows an availability level per instance type, so you know what can be deployed before you try.
- Waitlist: For GPUs that aren’t currently available, join a waitlist and get notified when they become available again.
- Quota increase requests: Request a quota increase directly in the console — no external inquiry needed. Requests can be made from the replica settings and the limits page at any time.
- Better GPU list: The GPU list is sorted by hourly price, with a refresh button to pull the latest stock.
- Reuse a configuration: Create a new container from an existing one’s settings — instance, image, scaling, and more. Only the configuration is carried over; no data is copied.
- Edit before creating: Change the name, endpoint URL, and environment variables before the new container is created.
- Redesigned billing pages: Charging, payment methods, and payment history are organized into tabs, and payment history can be exported as CSV.
- Redesigned usage & settlement views: The organization usage pages have been rebuilt with filters, summary cards, aggregate charts, and a transaction timeline.
- Charge breakdown: Settlement details now show how each amount was computed (unit price × hours × instance count).
- Reserved usage, separated: Reserved (prepaid) usage appears in charts and tables distinctly from on-demand usage, with its relationship to total spend explained.
- Actual prices everywhere: Lists and the create flow show the price you’ll actually be charged, and the create dialog includes a 10-minute billing estimate.
- Unified credit unit: Balance and pricing displays now use a single credit (C) unit.
- Deployment complete: Get notified when a container deployment finishes.
- Reserved device lifecycle: Receive notifications for key reserved-device moments such as start and upcoming expiry — start and expiry are also delivered by email.
- Billing alerts: Get notified about low balance, auto top-up failures, and overdue invoices so you can act before service is interrupted.
- Notification guide: See which notifications arrive in the console, on your desktop, or by email from the guide on the notifications page.
- Mark as read individually: Mark notifications as read one by one, with the read state reflected instantly.
- Organization switcher: The console header now shows which organization you are working in. An account belongs to one organization, so the switcher only appears if you hold more than one membership.
- Model pages, unified: Public model pages now share the console’s new design, and you can try models in the Playground without signing in. Model descriptions are available in Korean.
- Per-replica logs: The log sidebar is now organized by replica, showing running/stopped status, with per-replica viewing and bulk download.
- Volume resize: Increase the size of a persistent volume that’s already in use.
- Shared memory size: Set the shared memory (shm) size when creating a container.
- Inline API key issuance: Issue an API key without leaving the Air API page, with the endpoint address and pricing pinned on screen.
- Better dashboard charts: Container dashboard charts are easier to read, with new “since deployment” and “since creation” date range presets.
- Refined container create flow: The persistent volume choice now comes before GPU selection, and the data-loss warning was moved somewhere more visible.
Redesigned console, real-time notifications, and Persistent Volume
Redesigned console pagesWe’ve refreshed the console’s most-used pages for a cleaner, more consistent experience.- Container create: A rebuilt create flow — a clearer default mode, a card-based layout, and a configuration summary before you deploy. Start from a template with recommended items and an auto-filled name.
- Organization overview: Redesigned organization page that surfaces status and usage at a glance.
- Project overview: A reworked overview that brings your recent containers, recommended quick-start templates, and an Air API usage summary together on one page.
- Consistent navigation: Cleaner, header-free pages and unified breadcrumb navigation throughout the console.
- Instant delivery: Events such as deployment status changes and announcements appear in real time.
- Notification center: Browse, page through, and dismiss past notifications from a single view.
- Balance & usage alerts: Get notified about low balance, GPU usage charges, and payment failures so you can act before service is interrupted.
- Works with autoscaling: Attach a volume to deployments that use autoscaling and multiple replicas.
- New or existing volume on create: Choose a fresh volume or reuse an existing one when creating a deployment.
- Shared across deployments: A single volume can be attached to multiple deployments, and the picker shows which deployments are already using it.
- Edit while stopped: Enable, disable, or swap the attached volume while a deployment is stopped.
- Container tags: Organize and filter your containers with tags.
- Replica quota at start: Replica quota is now checked when you start a deployment, with an in-line dialog to adjust replicas or request more quota.
- Web Terminal replica selection: When a container has multiple replicas, choose which one to connect to, and the session auto-reconnects if a replica is replaced.
- SSH key scoping: Apply a registered SSH key to specific containers, and enable or disable it without removing it.
- Inline email verification: Enter a verification code directly during signup; disposable email domains are now blocked at registration.
- Better Air API usage page: More accurate spend totals, consistent filters, and model display names across charts and records.
- Persistent list views: Sorting, filters, search, and pagination are kept in the URL, so your view is restored when you navigate back from a detail page.
- New brand color: Updated the primary brand color across the console.
Web Terminal, AirCloud CLI, SSH access, and Shared Storage
AirCloud endpoint access and storage capabilities have been expanded. You can now connect to running containers from the browser or your local terminal, manage endpoints with the AirCloud CLI, use SSH-based development workflows, and share persistent storage across endpoints and replicas.Web TerminalBrowser-based terminal access is now available for custom container endpoints.- SSH-free terminal access: Open a terminal session through AirCloud’s exec-based connection without installing
sshdor configuring SSH keys. - Custom container support: Custom container endpoints now show a Terminal button on the endpoint overview page.
- Template environment access: AirCloud-provided templates such as Jupyter Notebook and Code already include terminal access through their web UI.
- Replica selection: For endpoints with multiple replicas, select the replica you want to connect to. If only one replica is running, AirCloud connects automatically.
- Simple session exit: Exit the terminal with
exitorCtrl-D. The terminal window closes automatically shortly after the session ends. - Connection retry guidance: If the terminal stays in the connecting state for a long time, try again from a new browser window or incognito mode.
aircloud-cli lets you manage AirCloud endpoints and access running containers from your local terminal.- CLI configuration: Configure the API base URL and API key with
aircloud config. - Authentication context: Use
aircloud whoamito verify your organization, project, user, and API key context. - Endpoint management: List, inspect, start, stop, scale, and patch endpoints from your terminal.
- Replica inspection: View live replica status for active endpoints.
- Endpoint logs: Retrieve endpoint logs, list log files, inspect per-replica logs, and fetch specific log ranges for debugging.
- Exec-based shell access: Use
aircloud exec <endpoint_id>to open an interactive shell inside a running container without SSH setup. - Replica pinning: Connect to a specific replica with
aircloud exec <endpoint_id> -r <replica_id>. - Custom shell command: Start a different shell or command with
aircloud exec <endpoint_id> -c "/bin/sh".
- CLI-based SSH access: Use
aircloud ssh <endpoint_id>to connect to supported containers from your local terminal. - SSH key injection: Register an SSH public key in the AirCloud console and inject it into selected endpoints.
- Template support: AirCloud-provided Jupyter Notebook and Code templates support SSH access out of the box.
- Custom container support: To use SSH with custom images, install
sshd, add injected public keys toauthorized_keys, and start the SSH server inside the container. - Standard SSH workflows: Use SSH-based workflows such as
scpfile transfer,ssh -Lport forwarding, and VS Code Remote SSH. - Replica-specific SSH: Pin an SSH session to a specific replica with
aircloud ssh <endpoint_id> -r <replica_id>. - Tunnel-only mode: Open only the SSH tunnel with
aircloud ssh <endpoint_id> --tunnel-onlyfor advanced workflows.
Shared StorageShared Storage is now available for persistent volumes across endpoints and replicas.
- One volume, multiple endpoints: Attach the same persistent volume to multiple endpoints within a project.
- Shared storage across replicas: Multiple replicas in the same endpoint can mount and access the same volume.
- Data independent from containers: Keep datasets, model checkpoints, logs, and output artifacts independent from the container lifecycle.
- Flexible volume attachment: Create a new persistent volume when creating an endpoint, or attach an existing volume to another endpoint.
- Read-write shared access: Shared volumes support read-write access, but applications should handle concurrent writes and file locking.
AirCloud CLI Reference
Learn how to install and use the AirCloud CLI.
Web Terminal and SSH Guide
Learn how to access running containers.
Shared Storage Guide
Learn how to share persistent volumes across endpoints and replicas.
Air API General Availability
Air API is now generally available. Access AirCloud’s AI models through an OpenAI-compatible interface with transparent per-token pricing.- OpenAI-compatible API: Use existing OpenAI SDKs and code to access AirCloud models with minimal changes.
- Public pricing: Per-token pricing is now published for all models with pay-as-you-go billing.
- Air API Playground: Test models directly in the browser without writing code.
Get started with Air API
Learn how to use Air API.
Browse Models
Explore all available models and pricing.
AirCloud Zero Release (RC)
AirCloud Zero is now available. Deploy Air Container and run AI inference workloads at a lower cost by leveraging crowdsourced GPU resources.- Lower-cost container deployment: Deploy Air Container at a lower cost than standard AirCloud by leveraging crowdsourced GPU resources.
- Built for inference workloads: Best suited for AI inference workloads where cost efficiency and flexibility matter most.
- Selectable at deployment time: Choose AirCloud Zero as the Cloud Type when creating an Air Container endpoint.
AirCloud Zero is currently in Release Candidate (RC). Some features and availability may be limited until General Availability (GA).
External API expansion and new model endpoints
The External API now covers full endpoint lifecycle management. Three new AI models are available on the platform.External API: New endpointsProgrammatic control over your deployments is now complete with five new endpoints:- List Endpoints: Retrieve a paginated list of all accessible endpoints with status and configuration details.
- List Replicas: View live replica status for any active endpoint.
- List/Get Log Files: Access endpoint log files and download their contents for debugging.
- Patch Endpoint: Update runtime settings (replica count, scaling config) for inactive endpoints.
- Get API Key Context: Verify your API key’s authentication scope and permissions.
API Reference
View the complete API documentation.
Browse Models
Explore all available models and pricing.
Jupyter Notebook, Web IDE, and Persistent Volume support
New development environments and storage capabilities for GPU workloads.Jupyter Notebook environmentGPU-powered Jupyter Notebook environments are now available as ready-to-use templates. Start developing immediately without any setup.- Pre-configured with TensorFlow, PyTorch, and other major libraries
- Full GPU resource access for experimentation and model development
- Browser-based access — no local installation required
- Built-in Jupyter AI integration
- Write and run code in the browser without a local development environment
- Container-based isolated workspace for each session
- Built-in development and debugging tools
- Code assistant integration included
- Store model checkpoints, logs, and data files across container restarts
- Data management independent of container lifecycle
- Ideal for long-running jobs and iterative experiments
External API initial release
Launched the AirCloud External API for programmatic endpoint management. Control your inference endpoints without leaving your terminal or CI/CD pipeline:- Get Endpoint Status: Check current status, configuration, and health of any endpoint.
- Start/Stop Endpoint: Start or stop endpoints on demand via API.
- Scale Replicas: Adjust replica count for active endpoints to match traffic requirements.
API Reference
Get started with the External API.
Air API beta launch, observability, dashboards, and UX upgrades
Air API launches in beta alongside major improvements to observability, project dashboards, and security.Air API (beta)An OpenAI-compatible inference API is now available in beta. Access AirCloud’s AI models with a single API key.- OpenAI SDK compatible — swap models without changing your code
- API key authentication with per-key endpoint access control
- Test models interactively in the Playground
- Log download and bulk export: Download logs for offline analysis with single or batch export.
- Advanced log filtering: Regex-based search and time-range filtering for faster troubleshooting.
- Raw log (JSON) view: Inspect unprocessed log data alongside the structured view.
- Scaling history: Track autoscale events, timeout-triggered changes, and error-driven replica adjustments over time.
- Project dashboards: Monitor status and usage at the project level with dedicated views.
- Organization usage breakdown: View usage and cost distribution by project, endpoint, and instance type.
- Extended request metrics: Cumulative request counts, job error rates, and other operational metrics.
- Transaction history filtering: Filter billing records by type and status.
- Korean language support: Full Korean UI now available.
- First-time user onboarding: Guided tutorial flow for new users to deploy their first endpoint.
- In-product feedback: Submit feedback directly from within the platform.
- OpenAI-compatible API keys: Issue API keys with OpenAI-compatible format. Control which endpoints each key can access.
- HTTPS endpoints: Enforced HTTPS for all inference endpoints.
- Path-based endpoint addressing: Moved from port-based to path-based routing with support for custom endpoint URLs.
AirCloud General Availability
AirCloud is now generally available. With this release, Air Container graduates from beta to GA. Deploy and operate your own container images on AirCloud’s GPU infrastructure.Air Container GA- Deploy custom container images directly on GPU clusters
- Autoscaling and scheduled scaling to match traffic patterns
- Full control over runtimes, dependencies, and service configuration
- Persistent storage integration for models and data
Get started with Air Container
Learn how to deploy containers.
- Time-based scaling: Schedule minimum replica counts or enable/disable autoscaling by time of day.
- Custom endpoint URLs: Use your own URL identifiers instead of system-generated IDs for cleaner integration.
- API Playground: Test and integrate APIs interactively through the built-in Playground without writing code.
- One-click cluster deployment: Reduce repetitive setup with templated, automated cluster deployments.
- Actionable error messages: Detailed error guidance on failure screens to cut mean-time-to-resolution.
- Secure inference traffic (HTTPS): End-to-end HTTPS for all inference requests.
Usage insights, autoscaling controls, and performance tuning
New visibility into costs, smarter autoscaling, and infrastructure-level performance hardening.- Usage & billing dashboards: Monthly and daily usage visualization to track credit consumption and spending trends at a glance.
- Autoscaling sensitivity controls: Configure autoscale responsiveness (Heavy / Normal / Light) to match your workload’s characteristics—from bursty inference to steady throughput.
- Deployment template reuse: Reuse existing deployment configurations to quickly spin up new endpoints with the same settings.
- Predictable scaling behavior: Adjusted scale-in/out policies to reduce variability under varying load patterns.
- Concurrent request handling improvements: Improved stability for large-file transfers and concurrent request handling under peak traffic.
Zero-downtime operations
Stability improvements so platform updates and scaling events don’t disrupt running workloads.- Zero-downtime updates: Minimized service disruption during platform updates.
- Request drop prevention: Improved handling during scale-in and updates to prevent in-flight request loss.
- Restart stability: Event processing pipelines now handle component restarts without data loss or ordering issues.
- HTTPS adoption: Began HTTPS rollout as the platform security baseline.
Billing and team collaboration
Introduced credit-based billing and multi-user collaboration so teams can manage costs and share resources.- Credits-based billing: Manage usage costs through a prepaid credit system with top-up and balance tracking.
- Team invites: Invite team members to your organization via email for collaborative access to shared resources.
- Reserved capacity: Distinguish between reserved (term-based) and on-demand resources for predictable cost planning.
- Billing automation: Automated usage collection and settlement processing on scheduled intervals.
Air Container beta launch and platform foundation upgrades
Air Container beta is now available. Deploy custom container images on GPU infrastructure for the first time on AirCloud.Air Container (beta)- Deploy custom container images on GPU clusters
- Basic autoscaling and monitoring support
- Container health checks with automatic recovery
- Custom metrics for user workloads: Collect and display per-container metrics (e.g., vLLM) with configurable metric targets.
- Improved health checks: Better readiness detection for containers with long boot times to prevent premature termination.
- Faster runtime environments: Lighter, faster packaging and deployment structures for reduced startup times.
- Offline-friendly deployment: Support for restricted-network environments with offline operation scenarios.
- Better model caching: More efficient HuggingFace model cache sharing and management for faster cold starts.
- Reliable event capture: Hardened autoscale and operational event capture for more stable scaling decisions.

