Percepta / Compute services

GPU compute.
Private AI.

GPU hosting for research, simulation and AI inference.

Rent capacity for your own workloads, reserve a server for ongoing work, or discuss hosting for your own language models. UK-hosted compute and on-premises AI systems for businesses and research teams.

Drawn live by your device.

GPU compute

GPU capacity for research, simulation and AI.

Access & deployment

Compute services.

Run your own software on rented GPUs, or ask us to scope a managed deployment. Hardware access and application management are separate services, with clear responsibilities and pricing.

01 / Project-based

On-demand compute

Short-term GPU access for experiments, model evaluation, rendering or a time-limited project.

  • Discuss the GPU and software environment
  • Agree storage, duration and total cost
  • Run your own workload
Request project capacity
02 / Ongoing workloads

Reserved capacity

A longer-term arrangement for recurring jobs or a service that needs an agreed allocation of compute.

  • Single- or multi-GPU configurations
  • Defined term and resource allocation
  • Support and maintenance agreed up front
Discuss a compute contract
03 / Scoped deployment

Private AI systems

A local model in a UK-hosted environment or a system on your premises, designed around your business requirements.

  • Model and hardware selection
  • Deployment and data-flow planning
  • Optional application integration by agreement
Explore private AI options

Already comfortable managing your own instance? Selected Percepta capacity is also offered through Vast.ai. Marketplace pricing, availability and terms apply; enquire directly if you need a particular configuration.

Businesses / Research / Engineering

What could you run?

You do not have to be an AI company to benefit from GPU compute. The important question is whether your software can use it effectively.

Research & experimentation

Model evaluation, parameter sweeps and reproducible training runs. Reserve capacity for a study or run independent experiments across GPUs.

Simulation & modelling

GPU-enabled numerical analysis and engineering solvers. The application, precision requirements and dataset determine whether RTX or H200 is appropriate.

LLM inference & adaptation

Run chat, coding and document-understanding models; evaluate alternatives or carry out supported adapter fine-tuning. Size the deployment for context length and concurrent requests.

Private business AI

Internal assistants and search over approved documents, using a hosted or on-premises model. Retrieval, access permissions and application integration are scoped together.

Computer vision

Train and evaluate detection models, test new datasets or run batch image analysis. Keep development compute separate from the systems deployed at your sites.

Rendering & batch processing

Render queues, visualisation and media-processing jobs. Match the GPU to the renderer, required features and software licence before reserving capacity.

For simulation and research: not every solver is GPU accelerated. Tell us the application, precision requirements and dataset size. We can discuss a representative trial before you commit to a longer allocation.

Model size / Concurrency / Precision

GPU configurations.

Choose around the workload: the model you need to serve, the number of simultaneous requests, or the solver you need to run. Exact hardware, availability and pricing are confirmed in your proposal.

RTX PRO 6000 Blackwell96 GB GDDR7 with ECC

Edition

CoolingActive double-flow-through

CoolingActive blower fan

CoolingPassive air-cooled; requires qualified ducted server airflow

Parts list

Workstation Edition
  1. Form factor
    Extended-height, dual-slot
    Nominal dimensions (height x length)
    5.4 x 12 inches; additional installation clearance required
  2. Cooling
    Active double-flow-through

    Allow suitable chassis clearance, airflow and system power for this edition.

  3. The extended-height, dual-slot cooler uses two flow-through airflow paths for the 600 W design.

  4. GPU architecture
    NVIDIA Blackwell / GB202
    CUDA cores
    24,064
    Tensor cores
    752 / fifth generation
    RT cores
    188 / fourth generation
  5. Graphics memory
    96 GB GDDR7 with ECC
    Memory interface
    512-bit
    Memory bandwidth
    1,792 GB/s
  6. Total board power
    600 W

    Choose the PSU for the complete system, not the GPU alone.

  7. Auxiliary power
    1 x PCIe CEM5 16-pin

    The GPU has a 600 W board-power rating and requires one PCIe CEM5 16-pin connection.

  8. Host interface
    PCIe 5.0 x16
    NVLink bridge
    Not supported
  9. Display connectors
    4 x DisplayPort 2.1b
    Display operation
    Direct display output; enabling MIG compute mode disables physical display outputs

Parts list

Max-Q Workstation Edition
  1. Form factor
    Full-height, dual-slot
    Nominal dimensions (height x length)
    4.4 x 10.5 inches; additional installation clearance required
  2. Cooling
    Active blower fan

    Allow additional room for the power cable and blower airflow.

  3. The 300 W active blower design prioritises power and space efficiency. Check slot spacing, airflow and total system power with your system manufacturer.

  4. GPU architecture
    NVIDIA Blackwell / GB202
    CUDA cores
    24,064
    Tensor cores
    752 / fifth generation
    RT cores
    188 / fourth generation
  5. Graphics memory
    96 GB GDDR7 with ECC
    Memory interface
    512-bit
    Memory bandwidth
    1,792 GB/s
  6. Total board power
    300 W

    The 300 W rating is for the GPU. Check PSU capacity for the complete system.

  7. Auxiliary power
    1 x PCIe CEM5 16-pin

    The card requires one PCIe CEM5 16-pin connection.

  8. Host interface
    PCIe 5.0 x16
    NVLink bridge
    Not supported
  9. Display connectors
    4 x DisplayPort 2.1b
    Display operation
    Direct display output; enabling MIG compute mode disables physical display outputs

Parts list

Server Edition
  1. Form factor
    Full-height full-length, dual-slot
    Nominal dimensions (height x length)
    4.4 x 10.5 inches; additional installation clearance required
  2. Cooling
    Passive air-cooled; requires qualified ducted server airflow

    Use a server with the forced airflow required for this passive card.

  3. GPU architecture
    NVIDIA Blackwell / GB202
    CUDA cores
    24,064
    Tensor cores
    752 / fifth generation
    RT cores
    188 / fourth generation
  4. Graphics memory
    96 GB GDDR7 with ECC
    Memory interface
    512-bit
    Memory bandwidth
    1,597 GB/s
  5. Total board power
    Up to 600 W; configurable, system and cable requirements apply
  6. Auxiliary power
    1 x PCIe CEM5 16-pin

    Board power is configurable up to 600 W; the supported setting depends on the board, firmware, cabling and host system.

  7. Host interface
    PCIe 5.0 x16
    NVLink bridge
    Not supported
  8. Display connectors
    4 x DisplayPort 2.1b
    Display operation
    Disabled in factory-default display-off mode; display-on requires supported configuration
PERCEPTASolutions Ltd / Edinburgh
Sheet01 / 0302 / 0303 / 03
TitleRTX PRO 6000 Blackwell Workstation EditionRTX PRO 6000 Max-Q Workstation EditionRTX PRO 6000 Blackwell Server Edition
96 GB
GDDR7 memory with ECC
1,792 GB/s
Memory bandwidth
600 W
Total board power
PCIe 5.0
x16 host interface
96 GB
GDDR7 memory with ECC
1,792 GB/s
Memory bandwidth
300 W
Total board power
PCIe 5.0
x16 host interface
96 GB
GDDR7 memory with ECC
1,597 GB/s
Memory bandwidth
Up to 600 W
Configurable board power
PCIe 5.0
x16 host interface
ViewExploded axonometric
Single GPU

RTX PRO 6000

96GBper Blackwell RTX PRO 6000 GPU / GDDR7 ECC

LLM hosting, a development environment or a single-GPU job.

LLM serving
Model weights, context cache and runtime must fit together. Quantisation, context length and concurrent requests determine capacity.
Other workloads
Model evaluation, supported adapter fine-tuning, computer vision training and GPU rendering.
Before booking
Validate the model, precision and software stack against your latency or job-time requirement.
Discuss a single-GPU setup
Multiple GPUs

RTX PRO 6000 servers

96GB / GPUSeparate memory on each GPU

More independent work, or a model distributed across GPUs.

More concurrent requests
Run multiple model replicas when each copy fits on one GPU. Route requests between workers.
Larger models & training
Use tensor or pipeline parallelism, or distributed training, where the model and framework support it.
Research & batch work
Run separate experiments, simulation jobs or render workers. Check data transfer and server topology before scaling one job.
Discuss a multi-GPU setup
AI & scientific compute

H200 servers

141GBper NVIDIA H200 GPU / HBM3e

Memory-heavy inference and supported scientific workloads.

LLM serving
More memory per GPU for weights and context cache. HBM3e bandwidth is relevant to memory-bound inference.
Simulation & numerical work
Double-precision (FP64) capability for compatible scientific software; choose by solver requirements, not AI performance figures.
Before booking
Confirm the H200 variant, GPU count, interconnect and application support. Test the intended workload.
Discuss H200 capacity

VRAM is not a model-size guarantee. Allow for model weights, the KV cache used by active requests and runtime overhead. Multiple GPUs need explicit software support to share a model; their memory is not automatically pooled.

Manufacturer specifications: RTX PRO 6000 / H200. These are configuration options, not a live availability list. CPU, system RAM, storage and networking are sized alongside the GPUs.

Language models / Private deployments

LLM inference
& model hosting.

Run open-weight models on UK-hosted GPUs or a system on your premises. Choose the model for your application, then size the compute for the context, response time and number of users it needs to support.

What needs to fit in GPU memory
Simultaneous requests
6
Context length
4

A model loading successfully is not the same as serving it well under load. Longer conversations and more simultaneous requests change the memory budget and response time.

Model selection

Models for your application.

Models evolve quickly. We help you choose from the latest available options based on your tasks, data, privacy requirements and budget.

A smaller model may meet your needs with less hardware. We test quality, response time and memory use on your workload, then confirm the model, licence and configuration before deployment.

Language & reasoning

Chat, analysis & internal assistants

Internal chat, summarisation and analysis, with reasoning suited to the task. Evaluate quality, languages and response time on your own tasks.

Code

Development tools & code assistance

Code generation, explanation and repair. Test against your languages and repository context; IDE integration and tool execution are separate application work.

Documents & images

Document understanding

Interpret text, tables and charts in document images. Validate extraction on representative pages, with document preparation and access controls included in the scope.

Retrieval

Embeddings & reranking

Generate vectors for semantic search, then rank retrieved documents for relevance. These complement a language model; ingestion, indexing and permissions need their own setup.

Serving engines

vLLM & SGLang.

Both support batching and prefix caching. We compare them on the exact model, GPU and request pattern, rather than assume one is faster.

vLLM

Continuous batching and paged attention help schedule requests and manage KV-cache memory. Model support covers generation and retrieval workloads.

SGLang

RadixAttention supports reuse of shared prompt prefixes. A candidate for repeated-context serving, with model-specific distributed inference options.

Validate before deployment. We agree model/version and licence, context and concurrency limits, then test response quality, time to first token and output speed on the chosen hardware. Deployment assistance and ongoing management are scoped separately from GPU rental.

Discuss LLM hosting
Private AI / Deployment options

UK hosted.
Or on your premises.

For businesses that do not want confidential prompts or documents routinely sent to a public AI service. We can discuss local-model hosting and on-premises systems, with the controls and responsibilities agreed before deployment.

Your business to a UK-hosted environmentWithin your business networkExplore where your AI could run
UK-hosted deployment. Your team, in your business, reaches the allocated hosted environment at our UK colocation facility over an agreed secure connection. The AI model, GPU compute and approved business data sit inside a data boundary; access, tenant isolation, storage, logging, backups and administrator permissions are agreed before confidential information is moved.On-premises deployment. Your team, the local network and your on-site system (the GPU server running the AI model, the approved business data, and their power, cooling and networking) are all within your premises. A line to the outside, for external tools, updates and remote support, needs separate consideration.
UK hostedYour premises
01Where the model runsUK colocation facilityNo server room required

Run a selected model on allocated compute at our UK colocation facility. This is remotely hosted computing, not an on-premises system.

Your on-site systemKeep inference on site

A local model can process prompts and documents within your premises.

02Where your data goesAgreed secure connection

The AI model, GPU compute and approved business data sit inside a data boundary.

Local network

External tools, updates and remote support need separate consideration.

03Who is responsibleDefine the data boundary

Agree access, tenant isolation, storage, logging, backups and administrator permissions before moving confidential information.

Plan the physical setup

We scope hardware, power, cooling, networking and handover with you. Installation and ongoing management are agreed separately.

Deployment design is agreed per project. A private model alone does not guarantee confidentiality or regulatory compliance.

Deployment scope

Data access & integration.

For an internal assistant, agree which documents it can retrieve and whose permissions it follows. For a hosted model, agree access controls, logging and retention. We scope these alongside the model, not as an assumption about where the server sits.

Operating experience / UK colocation

Our hosting
infrastructure.

We already host customer GPU workloads through Vast.ai and operate company-owned compute in UK colocation. You deal directly with Percepta about the hardware, deployment and support you need.

01

Location & data handling

We confirm the hosting location and data-handling requirements for your deployment. Storage, backups and any external services are included in that review.

02

Operating responsibilities

Agree who manages the operating system, models, updates, backups and monitoring. Renting compute does not automatically include application support or backup protection.

03

Availability by agreement

Capacity, maintenance windows, response times and any service-level commitment are defined in your contract. Marketplace hosting history is not a guarantee for every application.

Cost & commitment

How we quote.

For research budgets and business projects, the useful number is the cost of getting the work done. We scope the resources and term before you commit.

01

Compute

GPU model and count, CPU allocation and system memory.

02

Data

Working storage, retention, backups and any transfer costs.

03

Time

Project hours or a longer-term reserved allocation.

04

Support

Self-managed compute, deployment assistance or an agreed managed scope.

Have an existing cloud estimate? Bring it with your workload details. We can compare equivalent resources and scope; savings depend on utilisation, software, support and data transfer.

Before you book

A few practical details.

Can I rent for one project rather than a long contract?

Yes. Discuss short-term access with us or browse available instances on Vast.ai. Direct bookings are quoted and confirmed before the allocation starts; availability and minimum terms depend on the configuration.

Will my simulation or model run faster on a GPU?

Only if the software can use the GPU effectively. We need the application and version, GPU support, licence conditions, precision requirements and a representative workload. Some simulations are CPU-bound or require a different class of accelerator.

Does private AI mean nobody else can access my data?

No. A local model changes where inference runs, but access permissions, administrator access, logs, external tools, backups and support still matter. We agree the deployment boundary and data-handling arrangements with you; hosted and on-premises options have different responsibilities.

Can you build and install a system for our premises?

We can scope a local AI workstation or server around your use case. Hardware, software licensing, installation, network integration, training and ongoing support are confirmed separately in the proposal.

Are software licences, backups and managed support included?

Do not assume they are included in a GPU rental. Bring details of any commercial software or restricted datasets. We confirm permitted use, customer responsibilities and any additional storage, backup or support services before the work starts.

Talk directly to Percepta

Tell us what
you need to run.

A first conversation does not need a perfect specification. The application, deadline and budget are a useful starting point.

  • Your software or model and intended use
  • GPU, memory and storage needs, if known
  • Project duration, deadline and budget
  • Hosted or on-premises; any data restrictions

Please describe the requirement without sending confidential datasets, credentials or personal records.

Send a secure enquiry

All fields marked * are required.

Please do not include passwords, payment-card details or sensitive personal information.

We use the details you provide to review and respond to your enquiry and protect the form from abuse. See our privacy notice.

Your choice

Optional video settings

These demonstration videos are hosted by Google/YouTube, separately from Percepta's on-device camera analytics. Loading a player sends your IP address to Google/YouTube and may let it use cookies or similar storage for measurement and advertising.

No YouTube player connects until you allow YouTube and explicitly load a video. Allowing here does not load any videos. Each player needs its own Load action, then Play; nothing starts automatically.

Your choice stays in memory while you browse these public pages and resets on a full page reload. We do not save a preference cookie or use browser storage.

Turning videos off immediately removes all loaded players. It cannot erase information YouTube has already received or cookies it has already set.

YouTube videos: off.