What Valkan scans
The inventory and evidence produced by each connected integration.
This is the partner-facing inventory map. A connected integration reads the provider's control plane and selected audit/log data; it does not proxy traffic, decrypt content, or mutate provider resources.
The scan produces inventory and evidence. Findings are derived from that evidence, so a permission gap or disabled log is surfaced as a coverage warning rather than silently looking clean.
AWS
| Service / component | Inventory and evidence |
|---|---|
| IAM | IAM users, access keys (metadata only), MFA devices, and whether a user has signed in to the console (used only to label a user a person or a machine), owner tags, roles, groups and role bindings, including permissions inherited through groups and limited by Organizations service control policies. Also reachable principals, escalation paths, and federation / workload-identity relationships. Service-linked roles are skipped. Access analysis covers up to 200 roles and 200 users per account; anything beyond that is shown as a coverage gap. |
| Secrets Manager | Secret names and metadata, flagging secrets whose names suggest an AI-provider key and which roles can reach them; secret values are never read into Valkan. |
| S3 | Bucket names and creation dates, plus which identities' permissions reach each bucket. Bucket policies, ACLs and objects are not read. |
| Managed databases (RDS, Aurora, DocumentDB, Neptune, Redshift) | Database and cluster metadata: engine and version, endpoint host and port, region, VPC / subnets / security groups, public accessibility, encryption at rest, Multi-AZ, backup retention, status and tags. Where the control plane names one, the Secrets Manager secret holding the database's admin credentials is recorded as a link — the secret's VALUE is never read. Redshift's COPY / UNLOAD IAM roles are recorded as an access relationship. Database CONTENTS, schemas, table names, rows and query traffic are never read; this is control-plane metadata only. Aurora / DocumentDB / Neptune are listed once per cluster, not once per member instance. |
| EC2 | Running and stopped instances with type, VPC, subnet, private IPs, instance-metadata settings and tags (secret-looking tag keys are dropped), linked to the IAM role their instance profile grants, and to the users and roles whose policies let them start, stop or modify the instance (up to 100 instances per principal; read-only access is not shown, and IAM conditions are flagged, not evaluated). For each instance it also reads how the network reaches it: the security group rules open to the whole internet, whether its subnet routes to an internet gateway or NAT gateway, the internet-facing load balancers that have it as a target, and the Route 53 names that point at it or at its load balancer. Network ACLs, host firewalls and WAF rules are not evaluated, and a missing permission leaves exposure shown as unknown, not closed. The access graph draws the internet, any load balancer in front, and the ports each line opens. It also draws a line from a VM to each database it can connect to, when VPC peering, route tables and the database's security group all allow it. |
| EKS | Clusters with version, OIDC issuer, public-endpoint exposure and VPC, plus Pod Identity associations that map a Kubernetes service account to an IAM role. The Kubernetes API itself is not called. |
| Bedrock Agents | The draft version of each agent: instructions (first 200 characters), action groups, tools, Lambda executors, the guardrail and knowledge bases it references, tags, and agent-to-resource relationships. |
| AgentCore | Runtimes and gateways, in the regions where AgentCore is available (us-east-1 and us-west-2); other regions show it as unavailable. |
| Bedrock model access / logging | Model-invocation logging posture per region (off, or on to CloudWatch Logs or S3), plus coverage gaps where model access can't be read. |
| Bedrock usage metrics | Hourly per-model invocation, token, error, throttle and latency counts from CloudWatch, without caller identity. |
| SageMaker | All SageMaker endpoints, with status, configuration, production variants and the model image each runs. |
| Amazon Q Business | Q Business applications with their account / region metadata, data sources and web experiences. |
| Lambda | Functions that appear agent-related from environment-variable keys, layers, container images, or reserved concurrency. Values are redacted; only keys and detection signals are retained. |
| ECS | Services that appear agent-related from task images, environment-variable keys, or ECS Exec being enabled. Environment and secret values are redacted. |
| CloudTrail | AI-relevant control-plane activity for Bedrock and AgentCore from the lookback window (30 days by default). For the standard Bedrock Runtime calls (InvokeModel, streaming InvokeModel, Converse, and streaming Converse), Valkan keeps hourly caller-to-model aggregates and a source-IP candidate—not CloudTrail event/request IDs, prompts, responses, or request bodies. Agent lifecycle events are used only to detect unused agents. Secrets Manager GetSecretValue events (a standard management event) are read for who read which secret: Valkan keeps the caller, the secret, a count, and first and last read, never the secret value or the request. A read that names the secret only by name is matched to it through the account's secret list (secretsmanager:ListSecrets); a name that does not match is not drawn. RunInstances events (also standard management events) give who launched each VM inside the lookback window: the caller and the launch time, never the request. Reads of S3 objects are data events that are off by default and are not collected. Bedrock Mantle inference is not yet collected because it requires separately configured CloudTrail data-event capture. A denied, capped, or incomplete scan is shown as a coverage gap rather than no use. |
| Route 53 Resolver DNS | AI-provider and agent egress signals from Resolver query logs sent to CloudWatch Logs, attributed to the VM, Lambda function or ECS task that made the query. Lookups of database endpoints and of S3 bucket hostnames (<bucket>.s3.<region>.amazonaws.com) become observed workload-to-database and workload-to-bucket links, shown as probable: resolving a name does not prove any access. Path-style S3 addresses do not name the bucket in DNS and are not linked. Also logging coverage / blind spots per VPC. |
| AWS Organizations | Organization identity, management account, member-account roster, account hierarchy and service control policies when the connected principal can read them. |
The AWS scan covers the regions selected for the integration. Account-wide items (IAM, S3, Secrets Manager, access analysis) are read once per account, from the first selected region. When connected through an Organizations management account, each member account is scanned the same way.
Not covered on AWS today:
- DocumentDB Elastic clusters and Neptune Analytics graphs. Despite the names, these are separate AWS services with their own IAM prefixes (
docdb-elastic:,neptune-graph:) — therds:grants Valkan requests do not reach them, and Valkan does not request theirs. Standard DocumentDB and Neptune clusters ARE covered, because both answer on the RDS control plane. - Secrets Manager secrets outside the first selected region.
- MCP tool-call activity.
- DNS query logs delivered to S3 or Firehose, and DNS from Lambda functions that are not attached to a VPC.
Google Cloud
| Service / component | Inventory and evidence |
|---|---|
| IAM / service accounts | Service accounts, user-managed key metadata, the service accounts' IAM bindings on each project (and on the organization, for projects directly under one), the custom IAM role catalog, escalation paths, and workload-identity / federation / impersonation relationships. Reachable data is worked out for Cloud Storage buckets, Secret Manager secrets, Cloud SQL instances and KMS keys, from custom roles and from the built-in roles that grant data access, and an identity's broader abilities (for example Compute Engine or IAM) are summarized per service where no specific resource was found. People and groups named on a project, organization or service-account policy appear as identities, with the service accounts each can act as through an impersonation role on the account or its project; group membership is not read. Access from a service account in another project is flagged, and escalation paths cover impersonation chains, act-as-and-deploy, changing another account's access policy and creating a key for it (rated critical across projects). Private key material is never collected. The service-account key you connect Valkan with is listed like any other key, so it appears in the inventory as an API key. |
| IAM deny policies | Permission-dependent: policies attached to a scanned project and its ancestor folders and organization. Unconditional rules naming an identity or all principals reduce its supported data reach and service capabilities; a denied escalation permission is reported as blocked. Conditions, group membership and scoped service-account sets are not evaluated, and Owner/Editor wildcard access is not expanded. An unreadable policy produces a coverage gap and does not reduce access on a guess. Condition expressions are never stored. |
| BigQuery | Datasets in the selected regions (and the US / EU multi-regions on a selected continent), with location, labels and default encryption, plus each dataset's own access list: who can read, write or own it, and whether it is shared with all authenticated users. Tables, rows and queries are never read. Service accounts that can read a dataset are on its blast radius, whether the grant is on the dataset or on the project. |
| Pub/Sub | Topics, with labels, retention and encryption, plus who can publish to or subscribe from each one. Messages are never read. |
| Workload Identity Federation | Pools and their providers: the external issuer each one trusts, the names of the attributes it maps and checks, and the federated principals it issues. Condition values, SAML metadata and certificates are not stored. A provider for a shared issuer such as GitHub Actions with no attribute condition is flagged, because any of that issuer's customers can use it. |
| Change history | From Admin Activity audit logs (always on): who last changed each dataset, topic, KMS key, Cloud SQL instance, service account, workload identity pool and conversational agent, and the last ten changes. Request contents are not read. |
| Secret Manager | Secret metadata (rotation policy, labels, replication); secret values are never read into Valkan. |
| Cloud Storage (GCS) | Bucket names, location, storage class, uniform-access setting and labels. Bucket policies and objects are not read. |
| Vertex AI | Reasoning Engines (Agent Engine) agents, with metadata, instructions (first 200 characters), declared tools, project and location; plus Prediction endpoints (custom/AutoML model hosting) with their deployment status and request/response logging posture. Agents that show no recent use are flagged as inactive. |
| Vertex AI Agent Builder | Chat and search apps (Discovery Engine), the data stores they answer from, and the Dialogflow agent behind a chat app, and when each app was last used where Data Access audit logs are on. Indexed documents are never read. |
| Dialogflow | CX and ES agents, with each webhook's host and path, how it authenticates, and the names of the headers it sends. Header values, passwords, OAuth secrets and any token in a webhook URL are never stored; conversations are not read. A webhook that calls a collected Cloud Run service or Cloud Function is linked to it. When each agent was last used is read from Data Access audit logs where they are on. |
| Vertex AI usage | Per model and hour: calls, input and output tokens and latency from Cloud Monitoring, and — where Data Access audit logs are on — which service account or person made the calls. Request and response content is never read. Where calls were counted but no caller was recorded, Valkan says caller attribution is unavailable rather than showing none. A service account seen calling a model, with no Cloud Run service, function, agent or VM known to run as it, is listed as an inferred workload, grouped by account; a service account can be shared, so it is marked inferred. |
| Compute Engine | Every virtual machine in the selected projects and regions, with its status, machine type, zone, network placement and internal / external addresses, plus the service account it runs as and that account's access scopes. Stopped machines are included: they still hold their disks and their service account. Instance metadata is recorded as key names only — values, including start-up scripts, are never collected. Each machine also carries the firewall rules in its own project that let traffic in, and whether any opens it to the internet; Shared VPC host-project rules and organization firewall policies are not read. |
| Cloud SQL | Every instance in the selected projects and regions: engine, region, public or private address, authorized networks, high availability, backups and IAM database authentication, plus which service accounts, users and groups are IAM database users. Passwords, connection strings and data are never read, and password users are counted, not named. |
| Cloud KMS | Keys in the selected regions, their multi-region and global: purpose, protection level and rotation, plus who can use each key to encrypt or decrypt. Key material is never readable and is not requested. |
| Model Armor | Each project's Model Armor floor setting (and its parent folder's or organization's, where readable) and the templates in the selected regions. A Vertex AI agent is covered when an enforced floor includes Vertex AI; if the setting cannot be read the agent's guardrail status is unknown. Prompts and responses are never read. |
| Cloud Run | Services that appear agent-related from container images or environment-variable keys. Values are redacted. |
| Cloud Functions | Functions with a trigger or URL whose environment-variable keys look agent-related. Runtime values are redacted. |
| GKE | Clusters in every location, with version, workload identity pool, public-endpoint and authorized-network settings, and node count. Optionally, the Deployments, StatefulSets, DaemonSets and CronJobs inside each cluster: namespace, Kubernetes service account, images, privileged and host-network flags, and environment-variable names, with the Google service account each runs as where Workload Identity binds it. Secret values, Secret objects and pod logs are never read; a cluster reachable only privately is reported as not readable. |
| Cloud Logging | Audit events for Vertex AI (model calls and agent / endpoint changes), the Gemini API, Cloud Functions and Cloud Run from the lookback window (30 days by default). Only the method, service, caller, resource, severity, time and error code are kept; request and response bodies are not. Results are bounded and budget overruns are surfaced. |
| Cloud DNS | Queries to AI providers from DNS query logs, grouped by the VM or IP that made them (other queries are discarded), plus networks without DNS logging enabled. |
| Organization / folder context | The project's parent organization or folder and its name, when readable. With the optional Cloud Asset Inventory sweep, every project under the organization or folder is listed, each one Valkan cannot read is reported, and roles granted on folders count toward reach. |
A scan covers every active project the service account can see, except those you untick on the integration page. Changes apply from the next scan: an unticked project is no longer read, and the next completed scan removes what was found in it (resources, identities and findings) from the inventory and graph. Open cases close as out of scope because you excluded the project, without claiming the issue was fixed. These cases show “Out of scope (project excluded)” and do not count as resolved fixes. Tick it again and the following scan brings its data back; issues seen again recur on their existing cases. Cloud Run, Cloud Functions and Vertex AI are scanned in the regions selected for the integration (Vertex AI agents also in the global location); the other services are read project-wide. A service account may be valid while still lacking access to one project, API, or log stream; Valkan surfaces that as a permission gap.
Google Workspace
| Service / component | Inventory and evidence |
|---|---|
| Directory | Workspace users: email, display name, organizational unit, suspended state and last login. |
| OAuth grants | Third-party OAuth clients, granted scopes, affected users, known-tool matches, and the data stores implied by those scopes. |
| Token audit | Grant events from the Admin Reports API, used to date each OAuth grant. |
| Drive audit | File and folder access by third-party OAuth apps, sharing / ACL changes, and files or folders shared with service accounts or AI-related identities. Each record keeps the file's title, type and owner; content is not collected. |
| Admin audit | User create, update, delete and suspend actions carried out by AI apps or service accounts. |
Google Workspace does not inspect document contents, Gmail or Calendar access history, workflow internals inside third-party tools, or native Gemini activity. Admin role assignments are not read today. Drive audit events are available only on Business and Enterprise editions.
Microsoft 365 / Entra
| Service / component | Inventory and evidence |
|---|---|
| Directory | Entra users, groups, devices, directory role assignments (who holds an admin role), and app / service-principal registrations. |
| Application grants | Delegated (per-user consent) and application (tenant-wide) OAuth permission grants, kept visually distinct. |
| Sign-in activity | A bounded per-user rollup (last sign-in, sign-in count) from Entra sign-in logs — not one row per sign-in event. Requires Entra ID P1 or P2; below that, this is a coverage gap, not an empty result. |
| Directory audit | Administrative actions from the Entra audit log. Available on Entra ID Free. |
| Sign-in risk | Entra ID Protection risk detections. Requires Entra ID P1 or P2; shown as a coverage gap on lower tiers rather than an empty result. |
| SharePoint / OneDrive | Sites and drives, plus drive-root and individual-item sharing — direct, group, "anyone with the link," and app-only shares are each distinguishable for broadly-shared or app-shared items. Site-level sharing settings and personal OneDrive are not read; file contents are never read. |
| Teams | Installed Teams apps and their consented (resource-specific) permissions, per team. Chat and channel content is never read. |
| Copilot agent inventory (preview) | Copilot Studio and Microsoft 365 Copilot Agent Builder agents — orchestration, model, authentication, channels, and sharing. Reads via a separate Microsoft API (Power Platform Inventory) authorized by a directory role assignment, not a Graph permission. Preview: not yet verified working end-to-end — treat any result (or its absence) as unconfirmed until this note is removed. |
Microsoft 365 / Entra does not read mail, chat, or file contents, and does not collect Microsoft Purview unified audit data (a separate API and subscription flow, out of scope for this connector). Coverage depends on the tenant's Microsoft 365 / Entra plan — see Set up a Microsoft 365 / Entra account for what each tier unlocks. This is a separate connector from Azure — connecting one does not connect the other.
Endpoint and DNS inputs
These are separate from the cloud and SaaS connectors:
- Valkan Sensor telemetry records privacy-safe process, hostname, connection, and activity signals from enrolled devices — including per-destination byte volume (a sampled lower bound), which surfaces a large upload finding when one process sends ≥ 10 MiB to a single remote address and port in a minute. Devices also flag behaviour changes: a new destination or port, a volume spike, a change in timing pattern, off-hours activity, suspicious DNS and beaconing.
- AWS Route 53 and GCP Cloud DNS logging provide provider-side DNS evidence when logging is enabled.
What Valkan does not collect
- Credential, access-key, token, or secret values.
- Private keys.
- Document, message, email, or prompt contents beyond explicitly masked or privacy-safe metadata needed for classification. Drive file titles and owners are kept as metadata.
- Network payloads or decrypted traffic.
- Provider mutations. Connectors are read-only; the separate ticketing and alerting features are explicit operator-configured outbound actions.
Coverage can vary with provider permissions, API availability, region / project selection, audit-log retention, and plan tier. The scan status and findings feed show these gaps explicitly.