Establish project architecture and design decisions

This commit is contained in:
2026-07-11 16:27:59 +02:00
parent 3946b88d12
commit 5b4d52b201
11 changed files with 217 additions and 0 deletions
@@ -0,0 +1,38 @@
# ADR 0006: Use Pydantic for Domain Models
## Status
Accepted
## Context
The application manages structured and versioned entities such as meetings,
recordings, transcripts, participants and AI-generated artifacts.
These entities must be:
- type-safe
- validated at runtime
- serializable to JSON
- loadable from persisted JSON
- independent of the storage implementation
Python dataclasses provide a lightweight representation but do not provide
the required validation and serialization behavior without additional code.
## Decision
The project will use Pydantic models for its domain model.
Domain models shall inherit from a shared project base model.
Storage-specific behavior, database access and processing logic must not be
implemented inside the domain models.
## Consequences
- Invalid data can be rejected at module boundaries.
- Domain objects can be serialized to and restored from JSON.
- JSON schemas can be generated for interfaces and documentation.
- Pydantic becomes a core project dependency.
- Domain models remain independent of the eventual storage backend.
+50
View File
@@ -0,0 +1,50 @@
# ADR 0007: Recording Archive and Source Deletion Policy
## Status
Accepted
## Context
Uncompressed WAV recordings require substantial storage space. However, the
original recording must remain available until transcription, diarization and
quality verification have completed successfully.
The system requires a controlled mechanism for compressing recordings and
removing temporary source files without risking data loss.
## Decision
Recordings shall initially be captured in a processing-friendly lossless
format, normally WAV.
After transcription and diarization have completed, the recording may be
converted to an archive format.
The preferred archive formats are:
- FLAC for lossless archival
- Opus for storage-efficient speech archival
MP3 may be supported as an export format but is not the preferred internal
archive format.
The original WAV file may only be deleted after:
1. processing completed successfully,
2. all required transcript artifacts were persisted,
3. the archive file was created successfully,
4. archive integrity was verified,
5. explicit user confirmation was received.
Checksums shall be stored for both source and archive files where practical.
The processing format and the archive format are intentionally separated.
## Consequences
- Temporary WAV files may require significant short-term storage.
- Permanent storage consumption is reduced.
- Lossless and storage-efficient archive profiles can coexist.
- Source deletion remains deliberate and traceable.
- Failed or incomplete processing must never trigger automatic source deletion.
+104
View File
@@ -0,0 +1,104 @@
# ADR 0008: Meeting Storage Layout
## Status
Accepted
## Context
A meeting produces multiple related artifacts during its lifecycle, including recordings, transcripts, AI-generated analyses, exports and metadata.
The project requires a storage layout that is:
- human-readable
- portable
- independent of the storage backend
- suitable for long-term archival
- compatible with future database indexing
---
## Decision
Each meeting shall be stored in its own directory.
The directory represents the canonical container for all meeting artifacts.
Meeting directories shall be named using a stable Meeting ID rather than the meeting title.
Recommended format:
```text
YYYYMMDD_HHMMSS_<short_uuid>
```
Example:
```text
20260711_154215_9d7a4c51
```
The meeting title is stored exclusively inside `meeting.json`.
A meeting shall follow the structure:
```text
meeting/
│
├── meeting.json
│
├── recording/
│ ├── source.wav
│ └── archive.flac
│
├── transcript/
│ ├── canonical.json
│ ├── working_v1.json
│ └── working_v2.json
│
├── artifacts/
│ ├── summary.md
│ ├── executive_summary.md
│ ├── action_items.json
│ ├── decisions.json
│ ├── risks.json
│ └── knowledge.json
│
├── export/
│ ├── protocol.md
│ ├── protocol.pdf
│ └── protocol.docx
│
└── metadata/
├── recording.json
├── transcription.json
├── diarization.json
└── llm.json
```
The file system is considered the primary storage format.
Databases are optional secondary indexes.
File names describe their role rather than their file format.
Examples:
- `source.wav`
- `archive.flac`
- `canonical.json`
- `working_v1.json`
This allows processing and archive formats to evolve without changing the overall storage layout.
---
## Consequences
- Every meeting is self-contained.
- Meetings can be copied, archived and restored independently.
- Backup procedures remain simple.
- Future storage implementations remain compatible.
- Database implementations can be rebuilt from the meeting directories.
- Meeting titles may change without affecting directory names or references.
- The storage layout remains stable even if processing or archive formats evolve in the future.
View File
View File
+25
View File
@@ -0,0 +1,25 @@
"""Shared base classes for the Meeting Assistant domain model."""
from datetime import datetime, timezone
from uuid import UUID, uuid4
from pydantic import BaseModel, ConfigDict, Field
def utc_now() -> datetime:
"""Return the current time as a timezone-aware UTC datetime."""
return datetime.now(timezone.utc)
class DomainModel(BaseModel):
"""Base class for persistent domain entities."""
model_config = ConfigDict(
extra="forbid",
validate_assignment=True,
use_enum_values=False,
)
id: UUID = Field(default_factory=uuid4)
created_at: datetime = Field(default_factory=utc_now)
updated_at: datetime = Field(default_factory=utc_now)
View File
View File
View File
View File
View File