Conferences >2009 16th IEEE International ...

Combining multimodal and temporal contextual information for semantic video analysis

Download PDF
Download References
Request Permissions
Save to
Alerts

Abstract:

In this paper, a graphical modeling-based approach to semantic video analysis is presented for jointly realizing modality fusion and temporal context exploitation. Overal...Show More

Metadata

Abstract:

In this paper, a graphical modeling-based approach to semantic video analysis is presented for jointly realizing modality fusion and temporal context exploitation. Overall, the examined video sequence is initially segmented into shots and for every resulting shot appropriate color, motion and audio features are extracted. Then, Hidden Markov Models (HMMs) are employed for performing an initial association of each shot with the semantic classes that are of interest separately for every modality. Subsequently, an integrated Bayesian Network (BN) is introduced for simultaneously performing information fusion and temporal contextual knowledge exploitation, contrary to the usual practice of performing each task separately. The final outcome of the overall video analysis approach is the association of a semantic class with every shot. Experimental results as well as comparative evaluation from the application of the proposed approach in the domain of news broadcast video are presented.

Published in: 2009 16th IEEE International Conference on Image Processing (ICIP)

Date of Conference: 07-10 November 2009

Date Added to IEEE Xplore: 17 February 2010

ISBN Information:

ISSN Information:

DOI: 10.1109/ICIP.2009.5413673

Conference Location: Cairo, Egypt

Contents

1. INTRODUCTION

The rapid advances in hardware technology have led to a tremendous increase in the total amount of video content generated and distributed everyday. As a consequence, the need for efficient and advanced methodologies regarding video manipulation emerges as a challenging and imperative issue. To this end, several approaches have been proposed in the literature for tasks like search and organization of video content. More recently, the fundamental principle of processing the audio-visual information from a semantic-oriented perspective has been widely adopted, thus attempting to bridge the so called semantic gap [1] and efficiently capture the underlying semantics of the content.

References is not available for this document.

Combining multimodal and temporal contextual information for semantic video analysis

Abstract:

Metadata

Abstract:

ISSN Information:

1. INTRODUCTION

References

IEEE Account

Purchase Details

Profile Information

Need Help?

Combining multimodal and temporal contextual information for semantic video analysis

Alerts

Abstract:

Metadata

Abstract:

ISSN Information:

1. INTRODUCTION

Authors

Figures

References

Keywords

Metrics

Footnotes

References

IEEE Account

Purchase Details

Profile Information

Need Help?