A General Multimodal Teaching Video Generation System Based on Storyboard Prediction for Chinese Educational Corpora
DOI: https://doi.org/10.62517/jbdc.202601319
Author(s)
Ying Tian1, Yuhe Ding1, Wanyue Liu1,2,*
Affiliation(s)
1School of Artificial Intelligence and Big Data, Henan University of Technology, Zhengzhou, Henan, China
2iFLYTEK Co., Ltd., Hefei, Anhui, China
*Corresponding Author
Abstract
This paper analyzes how to build a video generation system with the help of RF-MTTextCNN model, and can translate the text content into a series of storyboards. In this system, six attributes can be predicted, including scene, difficulty and so on. The existence of these attributes can help the model generate storyboards, adjust the duration, and ensure that the animation rendering can be carried out smoothly. In this study, the system is tested with 4991 samples, and the three random seeds are introduced to find that the system's accuracy reaches 0.837, which is 0.061 higher than the existing model, and the macro-F1 value is 0.788, which is 0.196 higher than the existing model. Three test videos have H.264 video and AAC audio tracks, with a time difference of less than 0.02 seconds. The research results show that the system has strong engineering stability, good interpretability and strong controllability. However, the definition of weak labels is the same as that of rule experts, and the results are only consistent internally, but not consistent with the data marked by human beings.
Keywords
Teaching Video Generation; Storyboard Prediction; Rule Fusion; Multitask Learning; TextCNN; Audio-Video Synchronization
References
[1]Guo Jianwei. Adding a Beautiful Cover to Flash Animation. Computer Knowledge and Technology, 2016(3):96-97.
[2]Wang Bin. Batch Output for Unattended Multi-Video Rendering with Edius. Computer Knowledge and Technology, 2016(6):29-30.
[3]Singer U, Polyak A, Hayes T, et al. Make-A-Video: Text-to-Video Generation without Text-Video Data. arXiv:2209.14792, 2022.
[4]Kondratyuk D, Yu L, Gu X, et al. VideoPoet: A Large Language Model for Zero-Shot Video Generation. Proceedings of the 41st International Conference on Machine Learning, 2024:25105-25124.
[5]Yang Y, Fei Z, Xu D, et al. CogVideoX: Text-to-Video Diffusion Models with an Expert Transformer. arXiv:2408.06072, 2024.
[6]Ho J, Chan W, Saharia C, et al. Video Diffusion Models. arXiv:2204.03458, 2022.
[7]Caruana R. Multitask Learning. Machine Learning, 1997, 28:41-75.
[8]Kim Y. Convolutional Neural Networks for Sentence Classification. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, 2014:1746-1751.
[9]Vaswani A, Shazeer N, Parmar N, et al. Attention Is All You Need. Advances in Neural Information Processing Systems, 2017, 30.
[10]Ratner A, Bach S H, Ehrenberg H, et al. Snorkel: Rapid Training Data Creation with Weak Supervision. Proceedings of the VLDB Endowment, 2017, 11(3):269-282.
[11]Kim J, Kong J, Son J. Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech. Proceedings of the 38th International Conference on Machine Learning, 2021:5530-5540.
[12]Paszke A, Gross S, Massa F, et al. PyTorch: An Imperative Style, High-Performance Deep Learning Library. Advances in Neural Information Processing Systems, 2019, 32