gen_ai_hub.evaluations package¶
- class gen_ai_hub.evaluations.EvaluationClient(base_url: str, auth_url: str = None, client_id: str = None, client_secret: str = None, cert_str: str = None, key_str: str = None, cert_file_path: str = None, key_file_path: str = None, resource_group: str = None, aws_access_key_id: str = None, aws_secret_access_key: str = None, ai_core_client: AICoreV2Client = None, orchestration_url: str = None, input_object_store_secret_name: str = None, provider_name: str = 'aws')¶
Bases:
objectBase Client for the Evaluations service
- __init__(base_url: str, auth_url: str = None, client_id: str = None, client_secret: str = None, cert_str: str = None, key_str: str = None, cert_file_path: str = None, key_file_path: str = None, resource_group: str = None, aws_access_key_id: str = None, aws_secret_access_key: str = None, ai_core_client: AICoreV2Client = None, orchestration_url: str = None, input_object_store_secret_name: str = None, provider_name: str = 'aws')¶
EvaluationsClient root object to be used for Evaluations.
- Parameters:
base_url (str) – Base URL of the AI Core instance (must include /v2 suffix).
auth_url (str, optional) – Authentication URL used to retrieve access tokens.
client_id (str, optional) – OAuth client ID.
client_secret (str, optional) – OAuth client secret.
cert_str (str, optional) – X.509 certificate content as a string.
key_str (str, optional) – X.509 private key content as a string.
cert_file_path (str, optional) – File path to X.509 certificate.
key_file_path (str, optional) – File path to X.509 private key.
resource_group (str, optional) – Resource group name within the AI Core instance.
aws_access_key_id (str, optional) – AWS access key ID.
aws_secret_access_key (str, optional) – AWS secret access key.
ai_core_client (AICoreV2Client, optional) – Pre-configured AI Core client instance.
orchestration_url (str, optional) – Pre-existing orchestration deployment URL.
input_object_store_secret_name (str, optional) – Name of input object store secret.
provider_name (str, optional) – Hyperscaler provider name (e.g., “aws”).
- Raises:
ValueError – If required hyperscaler provider parameters are missing.
- create_or_update_object_store_secret(*, context, secret_body: dict, is_default: bool, result_key: str, attr_name: str, creator_mapping: dict, replace_existing: bool, result: dict)¶
- evaluate(evaluation_configs: List[EvaluationConfig]) List[EvaluationRun]¶
Main evaluate function to create the Evaluation job
- Parameters:
evaluation_configs (List[EvaluationConfig]) – A list of one or more of the EvaluationConfig objects
- Returns:
A list of EvaluationRun objects, one for each EvaluationConfig provided.
- Return type:
List[EvaluationRun]
- static from_env(profile_name: str = None, **kwargs)¶
Alternative way to create an EvaluationClient object.
Parameter resolution precedence: 1. Explicit keyword arguments 2. Environment variables 3. Configuration file 4. VCAP_SERVICES environment variable
- Parameters:
profile_name (str, optional) – Profile name defined in configuration.
kwargs – Additional parameters passed to constructor.
- Returns:
Configured EvaluationClient instance.
- Return type:
- get_system_supported_metrics() List[str]¶
helper method to get the list of all supported metric ids
- list_available_models()¶
Method to list all the available llm models
- resolve_orchestration_deployment_url() str¶
Resolves the orchestration deployment URL.
For non-default resource groups, creates a new deployment. For default resource group, attempts to discover existing deployment with the default config name using the orchestration service, or creates one if not found.
- Returns:
The orchestration deployment URL.
- Return type:
str
- setup(input_secret_body: dict | None = None, default_secret_body: dict | None = None, replace_existing: bool = False)¶
One time setup function which does object store secrets creation and orchestration deployment url creation if not provided.
- validate_secret_type(secret_type: str, creator_mapping: dict)¶
- class gen_ai_hub.evaluations.Dataset(source: str | Path | ArtifactSource)¶
Bases:
objectDataset object for the evaluations flow.
The Dataset class accepts various source types for evaluation datasets including local file paths (as strings or Path objects) or AI Core artifacts.
- Parameters:
source (Union[str, Path, ArtifactSource]) – Source of the dataset - can be a file path string, Path object, or ArtifactSource
Examples:
Using a Path object:
>>> Dataset(Path("data/sample.json"))
Using a string path:
>>> Dataset("data/sample.json")
Using an ArtifactSource with artifact dictionary:
>>> Dataset( ... ArtifactSource( ... artifact={ ... "id": "xyfz-rtyu-2456-ojns-yu6s", ... "name": "dataset-artifact", ... "url": "ai://default/eval_dataset" ... }, ... path="rootfolder/data.csv", ... file_type="csv" ... ) ... )
Using an ArtifactSource with artifact ID:
>>> Dataset( ... ArtifactSource( ... artifact="xyfz-rtyu-2456-ojns-yu6s", ... path="rootfolder/data.csv", ... file_type="csv" ... ) ... )
- __init__(source: str | Path | ArtifactSource)¶
Initialize a Dataset instance.
- Parameters:
source (Union[str, Path, ArtifactSource]) – Source of the dataset - can be a file path string, Path object, or ArtifactSource
- property file_type: str | None¶
Infer the file type from the source.
For ArtifactSource, returns the explicitly set file_type. For file paths, infers the type from the file extension.
- Returns:
File type (e.g., “json”, “jsonl”, “csv”) or None if cannot be determined
- Return type:
Optional[str]
- class gen_ai_hub.evaluations.EvaluationConfig(dataset_config: Dataset, metrics: List[MetricConfig], llm: LLMModelDetails | None = None, template: str | PromptTemplateSpec | TemplateRef | None = None, orchestration_registry_reference: str | None = None, template_variable_mapping: dict | None = None, test_row_count: int | None = -1, repetitions: int | None = 1, tags: dict | None = '{}', debug_mode: bool | None = False)¶
Bases:
objectDefines the evaluation configuration object for the Evaluations flow.
This class encapsulates all configuration parameters needed to run an evaluation job, including the model/template configuration, dataset, metrics, and execution settings.
At least one of the following must be provided:
llmandtemplatecombination (using orchestration_v2 models)orchestration_registry_reference(UUID of a registered orchestration configuration)
- Parameters:
dataset_config (Dataset) – Dataset configuration object specifying the evaluation dataset
metrics (List[MetricConfig]) – List of metric configurations for evaluation
llm (Optional[LLM]) – LLM configuration from orchestration_v2 (LLMModelDetails)
template (Optional[Union[str, PromptTemplateSpec, TemplateRef]]) – Prompt template as string, PromptTemplateSpec, or TemplateRef
orchestration_registry_reference (Optional[str]) – UUID of registered orchestration configuration
template_variable_mapping (Optional[dict]) – Variable mapping for the prompt template
test_row_count (Optional[int]) – Number of rows to sample from dataset (-1 for all rows), defaults to -1
repetitions (Optional[int]) – Number of times to repeat evaluation over the dataset, defaults to 1
tags (Optional[dict]) – User-defined metadata as key-value pairs, defaults to “{}”
debug_mode (Optional[bool]) – Enable debug logs in hyperscaler output path, defaults to False
Note
This module uses orchestration_v2 models directly.
Example using TemplateRef with ID:
>>> from gen_ai_hub.evaluations.models import EvaluationConfig, Dataset, MetricConfig >>> from gen_ai_hub.orchestration_v2.models.llm_model_details import LLMModelDetails as LLM >>> from gen_ai_hub.orchestration_v2.models.template_ref import TemplateRef, TemplateRefByID >>> config = EvaluationConfig( ... dataset_config=Dataset("data/test.jsonl"), ... metrics=[MetricConfig(name="accuracy")], ... llm=LLM(name="gpt-4", version="latest"), ... template=TemplateRef(template_ref=TemplateRefByID(id="template-id-here")), ... test_row_count=100 ... )
Example using TemplateRef with scenario/name/version:
>>> from gen_ai_hub.orchestration_v2.models.template_ref import TemplateRefByScenarioNameVersion >>> config = EvaluationConfig( ... dataset_config=Dataset("data/test.jsonl"), ... metrics=[MetricConfig(name="accuracy")], ... llm=LLM(name="gpt-4", version="latest", params={"temperature": 0.7}), ... template=TemplateRef(template_ref=TemplateRefByScenarioNameVersion( ... scenario="foundation-models", name="prompt1", version="1.0" ... )), ... test_row_count=100 ... )
- __init__(dataset_config: Dataset, metrics: List[MetricConfig], llm: LLMModelDetails | None = None, template: str | PromptTemplateSpec | TemplateRef | None = None, orchestration_registry_reference: str | None = None, template_variable_mapping: dict | None = None, test_row_count: int | None = -1, repetitions: int | None = 1, tags: dict | None = '{}', debug_mode: bool | None = False)¶
Initialize an EvaluationConfig instance.
- Parameters:
dataset_config (Dataset) – Dataset configuration object
metrics (List[MetricConfig]) – List of metric configurations
llm (Optional[LLM]) – LLM object from orchestration_v2 (LLMModelDetails), defaults to None
template (Optional[Union[str, PromptTemplateSpec, TemplateRef]]) – Prompt template (string, PromptTemplateSpec, or TemplateRef), defaults to None
orchestration_registry_reference (Optional[str]) – UUID of orchestration config, defaults to None
template_variable_mapping (Optional[dict]) – Variable mapping for prompt template, defaults to None
test_row_count (Optional[int]) – Number of dataset rows to sample (-1 for all), defaults to -1
repetitions (Optional[int]) – Number of evaluation repetitions (minimum: 1), defaults to 1
tags (Optional[dict]) – Key-value metadata pairs applied to all runs, defaults to “{}”
debug_mode (Optional[bool]) – Enable debug logging, defaults to False
- Raises:
ValueError – If neither (llm, template) nor orchestration_registry_reference is provided
- class gen_ai_hub.evaluations.MetricConfig(reference: MetricRef, variable_mapping: dict = None)¶
Bases:
objectDefines the metric config of the evaluation flow
- Parameters:
reference (MetricRef) – Provide the reference of metric to be evaluated, can be one of name,uuid(id), scenario/name/version
variable_mapping (Optional[dict]) – Any variable maping associated with the metric
- class gen_ai_hub.evaluations.MetricRef(scenario: str = None, name: str = None, version: str = None, id: str = None)¶
Bases:
objectRepresents a reference to a specific metric definition.
A metric can be identified in multiple ways: - By its UUID from metric management service (id) - By name (name) - By a combination of scenario, name, and version (scenario, name, version)
- __init__(scenario: str = None, name: str = None, version: str = None, id: str = None)¶
- class gen_ai_hub.evaluations.ArtifactSource(file_type: Literal['csv', 'json', 'jsonl'], artifact: str | Artifact, path: str | None = None)¶
Bases:
objectExtends the artifact object with the relative path user can provide inside to be used for EvaluationConfig Example Usage:
>>> ArtifactSource( artifact={ "id": "xyfz-rtyu-2456-ojns-yu6s", "name": "dataset-artifact", "url": "ai://default/eval_dataset" ... }, path= "rootfolder/data.csv, file_type="csv" ) >>> ArtifactSource( artifact="xyfz-rtyu-2456-ojns-yu6s", path="rootfolder/data.json, file_type="json" ) )
- __init__(file_type: Literal['csv', 'json', 'jsonl'], artifact: str | Artifact, path: str | None = None)¶
- Parameters:
artifact (Union[str,Artifact]) – Can just provide the artifact id as a string or the Artifact object of the AI_API_Client sdk.
path (Optional[str]) – Relative path within the artifact path provided and should point to a single file.
file_type (Literal["csv", "json", "jsonl"]) – One of the supported file_types
- class gen_ai_hub.evaluations.EvaluationRun(run_id: str, execution_id: str, ai_core_client: AICoreV2Client, configuration_id: str = None, artifact_id: str = None, resource_group: str = None, object_store_credentials: _AWSObjectStoreData = None, metrics_list: List[str] = None)¶
Bases:
objectRepresents an individual EvaluationRun object and its associated context.
- Parameters:
run_id (str) – Unique identifier for the evaluation run
execution_id (str) – ID of the AI Core execution
ai_core_client (AICoreV2Client) – AI Core client instance
configuration_id (str) – ID of the configuration, defaults to None
artifact_id (str) – ID of the artifact, defaults to None
resource_group (str) – Resource group name, defaults to None
object_store_credentials (_AWSObjectStoreData) – Object store credentials, defaults to None
metrics_list (List[str]) – List of metrics to evaluate, defaults to None
- __init__(run_id: str, execution_id: str, ai_core_client: AICoreV2Client, configuration_id: str = None, artifact_id: str = None, resource_group: str = None, object_store_credentials: _AWSObjectStoreData = None, metrics_list: List[str] = None)¶
- get_current_status()¶
Get the current status of the evaluation run.
- Returns:
Current status of the run
- Return type:
- Raises:
ValueError – If failed to retrieve the current status
- get_debug_info() ExecutionStatusDetails¶
Provide debug information when execution status is FAILED or DEAD.
- Returns:
Execution status details including failed pod information
- Return type:
ExecutionStatusDetails
- get_debug_logs()¶
Get the complete trace of execution logs.
- Returns:
List of log entries as dictionaries
- Return type:
list
- load_results_tables()¶
Download results from S3 and load the required table data.
- Returns:
Dictionary containing completions and metrics table data
- Return type:
dict
- Raises:
RuntimeError – If failed to download results
- results()¶
Get the results of the evaluation run.
- Returns:
Results object for accessing completion and metric results
- Return type:
- Raises:
ValueError – If execution is not completed
- set_cached_results_data(data)¶
Set the cached results data from the child results class.
- Parameters:
data (Any) – Results data to cache
- wait_for_completion(timeout: int | None = None)¶
Wait for the evaluation run to complete by polling status.
- Parameters:
timeout (Optional[int]) – Maximum time to wait in seconds, defaults to 3600 (1 hour)
- class gen_ai_hub.evaluations.Results(run: EvaluationRun)¶
Bases:
objectRepresents the Results handler for an EvaluationRun object.
This class provides methods to access completion results, metric results, and aggregated results for a specific evaluation run.
- Parameters:
run (EvaluationRun) – The parent EvaluationRun object
- __init__(run: EvaluationRun)¶
- aggregations()¶
Get the aggregated results for the run from the tracking service.
- Returns:
JSON response containing aggregated metric results
- Return type:
dict
- Raises:
ValueError – If error occurs while fetching aggregation results
- completions()¶
Get the completion results for the run.
- Returns:
DataFrame containing completion results for the run
- Return type:
pd.DataFrame
- Raises:
ValueError – If error occurs while fetching completions
- metrics()¶
Get the metric-level results for the run.
- Returns:
DataFrame containing metric results for the run
- Return type:
pd.DataFrame
- Raises:
ValueError – If error occurs while fetching metric results
Subpackages¶
- gen_ai_hub.evaluations.exceptions package
- gen_ai_hub.evaluations.models package
DatasetEvaluationConfigMetricConfigMetricRefArtifactSourceResultsEvaluationRun- Submodules
- gen_ai_hub.evaluations.utils package
- Submodules
- gen_ai_hub.evaluations.utils.aicore_utils module
generate_random_id()find_configuration_id_by_name()get_all_configurations()get_running_deployments_by_configuration_id()create_deployment_by_configuration_id()create_llm_orchestration_deployment_url()wait_for_target_status()read_data_from_artifact()build_s3_file_key()resolve_artifact_path()fetch_deployment_config()fetch_configuration_by_id()call_orchestration_service_with_v2_config()upload_file_to_aws_s3()upload_evaluation_dataset_data()register_aicore_artifact()register_aicore_configuration()register_aicore_execution()list_available_llm_models()fetch_orchestration_config_from_registry()resolve_metric_identifiers()resolve_metric_names()
- gen_ai_hub.evaluations.utils.config_data_utils module
- gen_ai_hub.evaluations.utils.file_utils module
- gen_ai_hub.evaluations.utils.gen_utils module
update_variable_mapping()get_accumulated_config_data()set_model_details_from_run_configs()create_model_versions_map_from_orch_configs()parse_model_filter_list()build_model_versions_map()create_model_versions_map_from_configuration_param_bindings()create_model_versions_map_from_custom_metric_config()select_model_details_randomly()update_test_orch_config()has_filter_key()get_filter_config()check_if_content_filter_provider_supported()remove_filter_metrics_if_provider_not_supported()create_custom_metric_name()validate_metric_name()count_user_prompts_from_template_list()get_template_list_from_orch_config()validate_prompts_in_templating_module()get_custom_metric_ids_from_input()check_if_metric_is_defined()is_value_in_json()validate_metrics()list_prompt_variables()extract_dataset_columns()get_prompt_variables_from_orch_config()get_grounding_config_from_orch_config()get_grounding_output_param_key()get_defaults()get_mapped_value_if_exists()validate_variable_mapping_of_prompts()validate_all_metrics_mapping()extract_metrics_variables()validate_individual_metrics()validate_variable_mapping_of_metrics()flatten_prompt_configuration()validate_individual_custom_metrics()handle_missing_dependent_variables_in_dataset()populate_dataset_data_if_data_missing()populate_dataset_data_if_single_schema_provided()handle_json_schema_match()validate_language_code_and_data_population()handle_language_match()populate_dataset_data_if_single_reference_provided()populate_dataset_data_if_individual_metric_reference_provided()handle_reference_missing_rows()update_artifact_dict()resolve_orchestration_config_v2()
- gen_ai_hub.evaluations.utils.language_match_utils module
- gen_ai_hub.evaluations.utils.metric_client_utils module
- gen_ai_hub.evaluations.utils.orch_config_utils module
validate_mandatory_modules()validate_orch_config_mandatory_modules()get_model_name()validate_model_name()validate_model_name_in_llm_module_config()get_prompt_templating_config()validate_template_ref_absent_in_config()get_template_key()validate_if_template_list_is_empty()validate_if_template_list_is_empty_in_templating_module_config()get_template_list_from_orch_config()validate_if_content_inside_template_is_empty_in_templating_module_config()validate_if_image_url_is_provided_in_content_type_inside_templating_module_config()validate_if_grounding_output_present_in_prompt_variables()validate_if_all_grounding_input_params_present_in_prompt_variables()validate_orchestration_params_from_evaluation_config()to_comparable()
- gen_ai_hub.evaluations.utils.oss_secret_utils module
- gen_ai_hub.evaluations.utils.validation_utils module
validate_filtered_models()fetch_and_validate_orchestration_config()validate_orchestration_url_across_configs()extract_deployment_id()validate_orchestration_url()validate_orchestration_configuration()validate_input_config()validate_variable_mapping_with_input_config()validate_config_data_collection()validate_merged_config_data()
- gen_ai_hub.evaluations.utils.aicore_utils module
- Submodules
Submodules¶
- gen_ai_hub.evaluations.client module
EvaluationClientEvaluationClient.__init__()EvaluationClient.from_env()EvaluationClient.validate_secret_type()EvaluationClient.create_or_update_object_store_secret()EvaluationClient.resolve_orchestration_deployment_url()EvaluationClient.setup()EvaluationClient.evaluate()EvaluationClient.list_available_models()EvaluationClient.get_system_supported_metrics()
- gen_ai_hub.evaluations.constants module
- gen_ai_hub.evaluations.credentials module